Depth policy that measures draft and verify cost per depth
Parent: Mac local LLMs: Speculative decoding and MTP · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Ollama: cost is an EWMA of target-forward wall time per visited depth, read by piecewise-linear interpolation, flat outside the sampled range. Acceptance is a per-position conditional EWMA. Expected committed tokens at depth N are 1 plus the running product of position acceptances summed to N.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Ollama: cost is an EWMA of target-forward wall time per visited depth, read by piecewise-linear interpolation, flat outside the sampled range. Acceptance is a per-position conditional EWMA. Expected committed tokens at depth N are 1 plus the running product of position acceptances summed to N. [source]
- Ollama: a depth is only costed when a round's depth equals the prior round's, so batch-shape transitions never enter the cost model; the controller seeds the shallowest unsampled depth to dwell there. [source]
- MTPLX: the policy had priced depth 2 below depth 3 although depth 2 ran the eager verifier and depth 3 the compiled fixed-M4 route; measuring per depth removed the mispricing. [source]
- MTPLX 2.11.2: measured costs, one-time compilation excluded, both depths re-probed; the 27B packs do not use this path. [source]
- Ollama: one stall-inflated cost sample could flip the comparison against plain decode and never heal once the controller stops parking at depth 0, hence the clamp. [source]
- Ollama: an unmeasured depth would always look best, so the search is capped at one past the acceptance frontier and at the drafter's limit. [source]
- Ollama: sessions with logprobs requests never draft and do not feed depth-0 cost. [source]
- Whether MTPLX's policy uses the same same-depth gate; the release notes do not say. [source]
- Ollama's cost EWMA weight is 0.3 (`costEWMAAlpha`) and one sample may move the estimate by at most 25 percent of its current value (`costClampFraction`). [source]
- The cost model is ready only after two distinct depths are sampled, and cost between samples is linearly interpolated. [source]
- Acceptance uses an EWMA weight of 0.1 per position, a position is trusted after 10 reaches (`acceptanceMinSamples`), and an under-sampled position inherits the deepest trusted rate, or 1 if none. [source]
- Only the surviving prefix is updated per round: position i is observed only if at least i-1 drafts were accepted. [source]
- The search limit is the acceptance frontier plus one, held to the drafter's limit, so depth climbs one position at a time. [source]
- Cost is recorded only when a round's draft depth matches the previous round's, and the controller drafts the shallowest unsampled depth to seed a clean sample. [source]
- The scheduled depth, cost curve, acceptance rates and probe cadence persist across requests in the speculation, and a new request starts at `depth.scheduled`. [source]
- A request with logprobs or top-logprobs keeps a speculation session only to maintain the draft cache and is permanently parked. [source]
- A round's draft is capped at remaining-1 tokens so the bonus token lands within the budget. [source]
- MTPLX 2.11.2 states the verify-cost mispricing and that `MTPLX_ADAPTIVE_VERIFY_COST_FEEDBACK=0` restores the configured prior. [source]
- MTPLX's trace report shows how many verify steps ran compiled and what a full round costs, and 2.12.0 adds an `mtp_pays` verdict comparing delivered tokens per second with a matched plain run. [source]
Children
- No children recorded.