<!-- llms-explorer concept facts · https://llms-explorer.com/tree/depth-policy-that-measures-draft-and-verify-cost/ · pack 2026-10-05 · ~1142 tokens -->

# Depth policy that measures draft and verify cost per depth

> Ollama: cost is an EWMA of target-forward wall time per visited depth, read by piecewise-linear interpolation, flat outside the sampled range. Acceptance is a per-position conditional EWMA. Expected committed tokens at depth N are 1 plus the running product of position acceptances summed to N.

Parent: [Mac local LLMs: Speculative decoding and MTP](https://llms-explorer.com/tree/mac-local-llms-speculative-decoding-and-mtp/) · 1 facets · 19 facts · page: https://llms-explorer.com/tree/depth-policy-that-measures-draft-and-verify-cost/

## Facts

- Ollama: cost is an EWMA of target-forward wall time per visited depth, read by piecewise-linear interpolation, flat outside the sampled range. Acceptance is a per-position conditional EWMA. Expected committed tokens at depth N are 1 plus the running product of position acceptances summed to N. — source: `asserted`
- Ollama: a depth is only costed when a round's depth equals the prior round's, so batch-shape transitions never enter the cost model; the controller seeds the shallowest unsampled depth to dwell there. — source: `asserted`
- MTPLX: the policy had priced depth 2 below depth 3 although depth 2 ran the eager verifier and depth 3 the compiled fixed-M4 route; measuring per depth removed the mispricing. — source: `asserted`
- MTPLX 2.11.2: measured costs, one-time compilation excluded, both depths re-probed; the 27B packs do not use this path. — source: `asserted`
- Ollama: one stall-inflated cost sample could flip the comparison against plain decode and never heal once the controller stops parking at depth 0, hence the clamp. — source: `asserted`
- Ollama: an unmeasured depth would always look best, so the search is capped at one past the acceptance frontier and at the drafter's limit. — source: `asserted`
- Ollama: sessions with logprobs requests never draft and do not feed depth-0 cost. — source: `asserted`
- Whether MTPLX's policy uses the same same-depth gate; the release notes do not say. — source: `asserted`
- Ollama's cost EWMA weight is 0.3 (`costEWMAAlpha`) and one sample may move the estimate by at most 25 percent of its current value (`costClampFraction`). — [source](https://raw.githubusercontent.com/ollama/ollama/main/mlxrunner/speculate_depth.go)
- The cost model is ready only after two distinct depths are sampled, and cost between samples is linearly interpolated. — [source](https://raw.githubusercontent.com/ollama/ollama/main/mlxrunner/speculate_depth.go)
- Acceptance uses an EWMA weight of 0.1 per position, a position is trusted after 10 reaches (`acceptanceMinSamples`), and an under-sampled position inherits the deepest trusted rate, or 1 if none. — [source](https://raw.githubusercontent.com/ollama/ollama/main/mlxrunner/speculate_depth.go)
- Only the surviving prefix is updated per round: position i is observed only if at least i-1 drafts were accepted. — [source](https://raw.githubusercontent.com/ollama/ollama/main/mlxrunner/speculate_depth.go)
- The search limit is the acceptance frontier plus one, held to the drafter's limit, so depth climbs one position at a time. — [source](https://raw.githubusercontent.com/ollama/ollama/main/mlxrunner/speculate_depth.go)
- Cost is recorded only when a round's draft depth matches the previous round's, and the controller drafts the shallowest unsampled depth to seed a clean sample. — [source](https://raw.githubusercontent.com/ollama/ollama/main/mlxrunner/speculate_depth.go)
- The scheduled depth, cost curve, acceptance rates and probe cadence persist across requests in the speculation, and a new request starts at `depth.scheduled`. — [source](https://raw.githubusercontent.com/ollama/ollama/main/mlxrunner/speculate.go)
- A request with logprobs or top-logprobs keeps a speculation session only to maintain the draft cache and is permanently parked. — [source](https://raw.githubusercontent.com/ollama/ollama/main/mlxrunner/speculate.go)
- A round's draft is capped at remaining-1 tokens so the bonus token lands within the budget. — [source](https://raw.githubusercontent.com/ollama/ollama/main/mlxrunner/speculate.go)
- MTPLX 2.11.2 states the verify-cost mispricing and that `MTPLX_ADAPTIVE_VERIFY_COST_FEEDBACK=0` restores the configured prior. — [source](https://mtplx.com/releases/2.11.2/)
- MTPLX's trace report shows how many verify steps ran compiled and what a full round costs, and 2.12.0 adds an `mtp_pays` verdict comparing delivered tokens per second with a matched plain run. — [source](https://mtplx.com/releases/2.11.2/)
