<!-- llms-explorer concept facts · https://llms-explorer.com/tree/ollama-dflash-block-diffusion-draft-model-in-mlx/ · pack 2026-10-05 · ~959 tokens -->

# Ollama DFlash block-diffusion draft model in mlxrunner

> Draft limit is blockSize-1. A round proposes n = min(maxTokens, blockSize-1) tokens after the unvalidated current token.

Parent: [Mac local LLMs: Ollama internals](https://llms-explorer.com/tree/mac-local-llms-ollama-internals/) · 1 facets · 21 facts · page: https://llms-explorer.com/tree/ollama-dflash-block-diffusion-draft-model-in-mlx/

## Facts

- Draft limit is blockSize-1. A round proposes n = min(maxTokens, blockSize-1) tokens after the unvalidated current token. — source: `asserted`
- The input is the anchor token plus n mask tokens, not the full trained block. The source says this is exact for causal layers and measured as free for bidirectional ones. — source: `asserted`
- Row i predicts the token at its own position, so the anchor row is dropped, and rows 1..n are sampled from one batched distribution. — source: `asserted`
- The draft KV holds only context rows from target hidden features, each derived from the feature at its own slot. The block is written then rewound, and accepted tokens return later as context. — source: `asserted`
- Pending feature rows buffer up to 256 tokens, then flush in one context-only forward. — source: `asserted`
- Sits beside the MTP drafter behind one `drafter` interface; the MTP path is the older one in this tree. — source: `asserted`
- A committed run that starts past the buffered frontier panics ("leaves a context gap"), since it would leave a hole in draft context. — source: `asserted`
- Every write path must rewind the previous block first, or new rows land after it and cache contents no longer match positions. — source: `asserted`
- Penalties see only committed history, not other rows of the block. — source: `asserted`
- Because context entries need no look-ahead, the cache trie keys need no pairing, unlike MTP's token-pair keys. — source: `asserted`
- Which DFlash checkpoints Ollama ships for MLX and their measured speedup; no fetched source gives it. — source: `asserted`
- `dflashDrafter.draftLimit()` returns blockSize-1. — [source](https://raw.githubusercontent.com/ollama/ollama/main/mlxrunner/dflash.go)
- `dflashPendingFlushTokens` is 256. — [source](https://raw.githubusercontent.com/ollama/ollama/main/mlxrunner/dflash.go)
- The drafter sends only the anchor and the rows being sampled; the source calls this exact for causal layers and "measured as free" for bidirectional ones. — [source](https://raw.githubusercontent.com/ollama/ollama/main/mlxrunner/dflash.go)
- The draft sees the block's mask tokens (`maskToken` from `BlockParams()`), with row 0 unused because row i predicts its own position. — [source](https://raw.githubusercontent.com/ollama/ollama/main/mlxrunner/dflash.go)
- Proposals are sampled from one batched distribution via `Sampler.Distribution` and `SampleDistribution`. — [source](https://raw.githubusercontent.com/ollama/ollama/main/mlxrunner/dflash.go)
- `commitBlock` rewinds the block out of the draft caches and every write path must call it first. — [source](https://raw.githubusercontent.com/ollama/ollama/main/mlxrunner/dflash.go)
- The draft cache keeps only context rows, and a context entry at slot S derives only from the target features at S. — [source](https://raw.githubusercontent.com/ollama/ollama/main/mlxrunner/dflash.go)
- Media items are ignored in `committed` because the target hidden at a slot already carries image content. — [source](https://raw.githubusercontent.com/ollama/ollama/main/mlxrunner/dflash.go)
- `propose` returns nil when no context has been committed yet. — [source](https://raw.githubusercontent.com/ollama/ollama/main/mlxrunner/dflash.go)
- `newSpeculation` chooses `newDFlashDrafter` when the draft implements `model.BlockDraft`. — [source](https://raw.githubusercontent.com/ollama/ollama/main/mlxrunner/speculate.go)
