<!-- llms-explorer concept facts · https://llms-explorer.com/tree/ollama-mlxrunner-speculative-depth-controller-co/ · pack 2026-10-05 · ~706 tokens -->

# Ollama mlxrunner speculative depth controller (committed tokens per second)

> `newSpeculation` builds the subsystem only when the checkpoint has a draft head; a block-draft model gets the DFlash drafter, anything else the MTP drafter.

Parent: [Mac local LLMs: Ollama internals](https://llms-explorer.com/tree/mac-local-llms-ollama-internals/) · 1 facets · 14 facts · page: https://llms-explorer.com/tree/ollama-mlxrunner-speculative-depth-controller-co/

## Facts

- `newSpeculation` builds the subsystem only when the checkpoint has a draft head; a block-draft model gets the DFlash drafter, anything else the MTP drafter. — source: `asserted`
- The drafter's `draftLimit` is passed to the controller, so a depth the drafter never produces is never scheduled. — source: `asserted`
- Depth 0 runs through an inner pipelined plain decoder (parked), which keeps the draft cache level. — source: `asserted`
- Replaced a heuristic schedule that grew toward a fixed cap (existing dossier). — source: `asserted`
- A drafted round that resumes from parking drains the inner decoder's in-flight sample as the current token and attributes no cost to that round. — source: `asserted`
- A stall can inflate a cost sample severalfold, so samples are clamped (see sibling dossier). — source: `asserted`
- No source gives measured tokens per second gained by the controller against a fixed depth on a Mac. — source: `asserted`
- `newSpeculation` returns nil when the model has no draft head, and then requests decode plainly. — [source](https://raw.githubusercontent.com/ollama/ollama/main/mlxrunner/speculate.go)
- The speculation picks the DFlash drafter when the draft model implements `model.BlockDraft`, otherwise the MTP drafter. — [source](https://raw.githubusercontent.com/ollama/ollama/main/mlxrunner/speculate.go)
- The controller's `drafterLimit` is set from the drafter's `draftLimit()`, which is blockSize-1 for DFlash and 0 (unbounded) for MTP. — [source](https://raw.githubusercontent.com/ollama/ollama/main/mlxrunner/dflash.go)
- A parked request runs an inner pipelined decoder and counts each plain token as a depth-0 round. — [source](https://raw.githubusercontent.com/ollama/ollama/main/mlxrunner/speculate.go)
- Resuming drafting from a parked stretch emits the inner decoder's sampled-but-unforwarded token as the current token and records no cost for that round. — [source](https://raw.githubusercontent.com/ollama/ollama/main/mlxrunner/speculate.go)
- Each request opens with `limit = depth.scheduled` so a new request starts at the proven depth. — [source](https://raw.githubusercontent.com/ollama/ollama/main/mlxrunner/speculate.go)
- Round wall time runs from one round start to the next and spans the next emit's sync. — [source](https://raw.githubusercontent.com/ollama/ollama/main/mlxrunner/speculate.go)
