Ollama mlxrunner speculative depth controller (committed tokens per second)
Parent: Mac local LLMs: Ollama internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
`newSpeculation` builds the subsystem only when the checkpoint has a draft head; a block-draft model gets the DFlash drafter, anything else the MTP drafter.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- `newSpeculation` builds the subsystem only when the checkpoint has a draft head; a block-draft model gets the DFlash drafter, anything else the MTP drafter. [source]
- The drafter's `draftLimit` is passed to the controller, so a depth the drafter never produces is never scheduled. [source]
- Depth 0 runs through an inner pipelined plain decoder (parked), which keeps the draft cache level. [source]
- Replaced a heuristic schedule that grew toward a fixed cap (existing dossier). [source]
- A drafted round that resumes from parking drains the inner decoder's in-flight sample as the current token and attributes no cost to that round. [source]
- A stall can inflate a cost sample severalfold, so samples are clamped (see sibling dossier). [source]
- No source gives measured tokens per second gained by the controller against a fixed depth on a Mac. [source]
- `newSpeculation` returns nil when the model has no draft head, and then requests decode plainly. [source]
- The speculation picks the DFlash drafter when the draft model implements `model.BlockDraft`, otherwise the MTP drafter. [source]
- The controller's `drafterLimit` is set from the drafter's `draftLimit()`, which is blockSize-1 for DFlash and 0 (unbounded) for MTP. [source]
- A parked request runs an inner pipelined decoder and counts each plain token as a depth-0 round. [source]
- Resuming drafting from a parked stretch emits the inner decoder's sampled-but-unforwarded token as the current token and records no cost for that round. [source]
- Each request opens with `limit = depth.scheduled` so a new request starts at the proven depth. [source]
- Round wall time runs from one round start to the next and spans the next emit's sync. [source]
Children
- No children recorded.