<!-- llms-explorer concept facts · https://llms-explorer.com/tree/acceptance-drift-of-a-base-mtp-head-on-a-distill/ · pack 2026-10-05 · ~934 tokens -->

# Acceptance drift of a base MTP head on a distilled trunk

> The head conditions on the trunk's hidden state, so any trunk shift lowers acceptance without any error being raised.

Parent: [Mac local LLMs: Speculative decoding and MTP](https://llms-explorer.com/tree/mac-local-llms-speculative-decoding-and-mtp/) · 2 facets · 18 facts · page: https://llms-explorer.com/tree/acceptance-drift-of-a-base-mtp-head-on-a-distill/

## Facts

- The head conditions on the trunk's hidden state, so any trunk shift lowers acceptance without any error being raised. — source: `asserted`
- Acceptance falls fastest at deeper positions, because each position is conditional on the whole earlier prefix surviving. — source: `asserted`
- The runtime reacts by choosing a shallower depth (MTPLX tune, Ollama depth controller), not by retraining the head. — source: `asserted`
- 2.12.0 (23 Sep 2026) is the first release to publish measured per-position acceptance for a grafted head (MiMo) and an overall figure for Bonsai. — source: `asserted`
- A head with wrongly shifted norm weights accepts about 0 percent, which is a different failure from drift (the right token sat at median rank 247,513 of 248,320). — source: `asserted`
- Calibration on too few prompts can flip a grafted head to a wrong setting; Forge now needs at least 4 prompts and 48 draft rounds per option. — source: `asserted`
- Bonsai's 75 percent overall acceptance sits next to MiMo's 56.5 percent at position 2. The sources do not say whether the two figures are comparable (different prompts, depth, thinking setting). — source: `asserted`
- MiMo's acceptance against the Qwen3.5-9B base on the same prompt (no baseline published). — source: `asserted`
- Whether acceptance at depth 1 for MiMo (78.5 percent) is high enough that depth 1 would beat the shipped depth 2. — source: `asserted`
- The Qwen3.5-9B draft head accepts 78.5 percent of drafts at the first position and 56.5 percent at the second on MiMo V2.6 Qwen 9B, measured on a long game prompt with thinking on. — [source](https://mtplx.com/releases/2.12.0/)
- The MTPLX notes state the MiMo head "accepts fewer drafts on MiMo" because it was trained on Qwen3.5-9B, and keep depth 2 as the default as on Qwen 9B. — [source](https://mtplx.com/releases/2.12.0/)
- The Qwen3.8-27B head accepts 75 percent of its drafts on Ternary Bonsai 2 27B. — [source](https://mtplx.com/releases/2.12.0/)
- With the new kernels, Bonsai at depth 3 was 14 percent slower than depth 1 over a 3,000-token answer, which is why Bonsai ships at depth 1. — [source](https://mtplx.com/releases/2.12.0/)
- A draft head imported from a raw checkpoint without the +1.0 norm correction loaded without error and accepted about 0 percent of drafts, with the right token at median rank 247,513 of 248,320. — [source](https://mtplx.com/releases/2.12.0/)
- Forge switches a head's norm setting only with at least 4 prompts, 48 draft rounds per option and a lead of 0.25 accepted tokens per round; otherwise the declared setting stays and the runtime file records the calibration as inconclusive. — [source](https://mtplx.com/releases/2.12.0/)
- The MiMo pack's speed was "not measured yet" at 2.12.0, with the same per-forward work as Qwen 3.5 9B but fewer accepted tokens. — [source](https://mtplx.com/releases/2.12.0/)
- Inference: a distilled trunk's drift is visible first at position 2 and later, so a per-position acceptance model is the right detector. — source: `asserted`

## Corrections and disagreements

- CONTRADICTS: grafting-a-base-model-s-mtp-head-onto-a-fine-tun.md Open question "draft acceptance of the MiMo pack": the 2.12.0 release notes do publish acceptance (78.5 and 56.5 percent); only speed is unmeasured. — source: `asserted`
