Acceptance drift of a base MTP head on a distilled trunk
Parent: Mac local LLMs: Speculative decoding and MTP · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
The head conditions on the trunk's hidden state, so any trunk shift lowers acceptance without any error being raised.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- The head conditions on the trunk's hidden state, so any trunk shift lowers acceptance without any error being raised. [source]
- Acceptance falls fastest at deeper positions, because each position is conditional on the whole earlier prefix surviving. [source]
- The runtime reacts by choosing a shallower depth (MTPLX tune, Ollama depth controller), not by retraining the head. [source]
- 2.12.0 (23 Sep 2026) is the first release to publish measured per-position acceptance for a grafted head (MiMo) and an overall figure for Bonsai. [source]
- A head with wrongly shifted norm weights accepts about 0 percent, which is a different failure from drift (the right token sat at median rank 247,513 of 248,320). [source]
- Calibration on too few prompts can flip a grafted head to a wrong setting; Forge now needs at least 4 prompts and 48 draft rounds per option. [source]
- Bonsai's 75 percent overall acceptance sits next to MiMo's 56.5 percent at position 2. The sources do not say whether the two figures are comparable (different prompts, depth, thinking setting). [source]
- MiMo's acceptance against the Qwen3.5-9B base on the same prompt (no baseline published). [source]
- Whether acceptance at depth 1 for MiMo (78.5 percent) is high enough that depth 1 would beat the shipped depth 2. [source]
- The Qwen3.5-9B draft head accepts 78.5 percent of drafts at the first position and 56.5 percent at the second on MiMo V2.6 Qwen 9B, measured on a long game prompt with thinking on. [source]
- The MTPLX notes state the MiMo head "accepts fewer drafts on MiMo" because it was trained on Qwen3.5-9B, and keep depth 2 as the default as on Qwen 9B. [source]
- The Qwen3.8-27B head accepts 75 percent of its drafts on Ternary Bonsai 2 27B. [source]
- With the new kernels, Bonsai at depth 3 was 14 percent slower than depth 1 over a 3,000-token answer, which is why Bonsai ships at depth 1. [source]
- A draft head imported from a raw checkpoint without the +1.0 norm correction loaded without error and accepted about 0 percent of drafts, with the right token at median rank 247,513 of 248,320. [source]
- Forge switches a head's norm setting only with at least 4 prompts, 48 draft rounds per option and a lead of 0.25 accepted tokens per round; otherwise the declared setting stays and the runtime file records the calibration as inconclusive. [source]
- The MiMo pack's speed was "not measured yet" at 2.12.0, with the same per-forward work as Qwen 3.5 9B but fewer accepted tokens. [source]
- Inference: a distilled trunk's drift is visible first at position 2 and later, so a per-position acceptance model is the right detector. [source]
Corrections and disagreements
- CONTRADICTS: grafting-a-base-model-s-mtp-head-onto-a-fine-tun.md Open question "draft acceptance of the MiMo pack": the 2.12.0 release notes do publish acceptance (78.5 and 56.5 percent); only speed is unmeasured. [source]
Children
- No children recorded.