Grafting a base model's MTP head onto a fine-tuned trunk (MiMo-V2.6 distill)
Parent: Mac local LLMs: Speculative decoding and MTP · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
The catalog maintainer builds the pack from measured evidence: quantized trunk, the fine-tune's own BF16 vision tower, and the base model's draft head. A user cannot do this by attaching a sidecar, per the README.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- The catalog maintainer builds the pack from measured evidence: quantized trunk, the fine-tune's own BF16 vision tower, and the base model's draft head. A user cannot do this by attaching a sidecar, per the README. [source]
- The head stays valid only as far as the fine-tuned trunk's hidden states stay close to the base trunk's. MTPLX checks this by measuring acceptance and by KL against the fine-tune's BF16 reference, not by comparing shapes. [source]
- Draft depth is a per-pack default set from evidence: depth 2 for MiMo, depth 1 for Bonsai. [source]
- 2.11.2: Forge extracts MiMo-layout heads and binds MiMo's output head (PR 442). [source]
- 2.12.0 (23 Sep 2026): the MiMo V2.6 Qwen 9B pack ships, 8.70 GB at 6-bit. The same release adds Bonsai 2 27B with the Qwen3.8-27B head. [source]
- 2.12.0 fix: a model named MiMo but built on the Qwen architecture was served with MiMo family settings (depth 1 only, reasoning off); family now comes from model type and architecture. [source]
- 2.12.0 fix: Forge keeps a grafted head's declared norm setting unless calibration evidence is clear. [source]
- Name-based family matching misrouted the Qwen-architecture MiMo distill to a different family's settings. [source]
- A calibration on one prompt and six tokens with zero accepted drafts once flipped a grafted head to `pre_norm`. [source]
- An output head can be left random: PR 442 notes MiMo's output head was bound instead of random. [source]
- The MiMo pack's decode speed was unmeasured at release; only memory peak and KL fidelity are published. [source]
- The README says MTPLX does not support attaching a separately supplied sidecar to an arbitrary trunk, yet the catalog ships two packs that carry a base model's head on a different trunk. These are consistent only if the maintainer's measured pack counts as the "complete model with matching MTP weights" the README allows. No source states that reading. [source]
- Draft acceptance and decode speed of the MiMo pack against plain Qwen 3.5 9B: the README, FAQ and releases page say speed was not measured. [source]
- How far a distillation can drift before a base head's acceptance falls below a no-draft baseline. [source]
- Xiaomi's MiMo-V2.6-Distill-Qwen-9B is a Qwen3.5-9B fine-tune for coding and agents under the MIT license, and its MTPLX pack carries "Xiaomi's BF16 vision tower and the Qwen3.5-9B draft head (the checkpoint ships none)". [source]
- The MiMo pack `Youssofal/MiMo-V2.6-Qwen-9B-MTPLX-Optimized-Speed` is 8,695,116,595 bytes with a 6-bit group-64 body, uses served id `mtplx-mimo-v26-qwen-9b-optimized-speed` and the Qwen 3.5 sampler contract (0.6, 0.95, 20) with depth 2 by default. [source]
- Against Xiaomi's BF16 checkpoint the MiMo pack measures KL 0.0054 and 97.3 percent top-1 agreement over 19,265 tokens. [source]
- The README catalog lists the MiMo pack as fitting 16 GB and up with a 10.0 GiB peak and the preset "Sustained, depth 2". [source]
- On a 16 GB Mac the MiMo pack plans a 20,480-token window, and the FAQ says its speed on MTPLX "has not been measured yet". [source]
- Xiaomi's model card reports 44.6 on SWE Pro against 32.0 for Qwen 3.5 9B, as relayed by the MTPLX FAQ. [source]
- The app and CLI list MiMo second on 16 to 31 GB Macs, after Ternary Bonsai 2 27B. [source]
- A 2.12.0 fix says the MiMo family had been matched by folder name, so Xiaomi's Qwen-architecture distills named MiMo were served with MiMo settings (draft depth 1 only, reasoning off), and the family is now read from the checkpoint's model type and architecture so the distill keeps Qwen 3.5 settings. [source]
- 2.11.2 notes say MiMo "reaches tune", Forge takes verification depths from the tune policy, and MiMo layouts are extracted by PR 442. [source]
- The Ternary Bonsai 2 27B pack carries Prism ML's weights byte for byte plus the Qwen3.8-27B draft head, 8.85 GB in total, with the draft head on by default at depth 1. [source]
- With the grafted head and two new kernels, Bonsai 2 27B decodes 64.4 tok/s at a 4K prompt and 57.1 at 16K on an M5 Max, against 52.6 and 51.0 for the 4-bit dense 27B in the same session, in about half the memory (11.4 GB against 23.9 GB peak). [source]
- Bonsai 2's measured mean KL to Prism ML's float32 reference is 2.92e-6 with MTPLX's ternary kernel against 3.02e-6 for stock, over 1,630 positions, with identical greedy output. [source]
- When tuning is skipped, the installed pack's own depth applies, which is why Bonsai starts at depth 1 instead of 2. [source]
- A Medium video review ran the MiMo V2.6 Qwen 9B distill on a 16 GB M1 MacBook Pro through the MLX Core app at about 12 tok/s on a short answer, which is a plain MLX run and not MTPLX drafting. [source]
Corrections and disagreements
- CONTRADICTS: training-an-mtp-adapter-for-a-model-with-no-nati.md Mechanism line "a fine-tune of a base model (MiMo-V2.6 distilled from Qwen3.5 9B) ships with the base model's MTP head". The changelog says the checkpoint ships none and MTPLX adds the Qwen3.5-9B draft head to its pack. [source]
Children
- No children recorded.