Training an MTP adapter for a model with no native head
Parent: Mac local LLMs: Speculative decoding and MTP · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Extraction: Forge copies an embedded MTP block into an `mtp.safetensors` sidecar, including heads stored under other prefixes (GLM-4 MoE, GLM-5.3-Flash, DeepSeek-V3.2, MiMo layouts), and keeps the vision tower.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Extraction: Forge copies an embedded MTP block into an `mtp.safetensors` sidecar, including heads stored under other prefixes (GLM-4 MoE, GLM-5.3-Flash, DeepSeek-V3.2, MiMo layouts), and keeps the vision tower. [source]
- Repair: it decides the norm convention once per tensor set (see the norm-convention dossier) and restores fused MoE experts. [source]
- Calibration: it probes wiring choices (hidden-state variant, history mode, depth) over sampled prompts and keeps the declared setting unless the evidence is clear. [source]
- Verification: it measures plain against speculative decoding and refuses a pack that is slower or not exact. [source]
- Grafting: a fine-tune of a base model (MiMo-V2.6 distilled from Qwen3.5 9B) ships with the base model's MTP head. [source]
- The earliest changelog entry for Forge says "calibrate and train the MTP adapter". [source]
- A later docs sweep removed unsupported MTP-sidecar graft guidance (issue 218, seeded by issue 215's "phantom graft script"). [source]
- PR 489 made Forge require MTP weight evidence and refuse AR-only trunks. [source]
- 2.12.0: MiMo V2.6 Qwen 9B joins with the Qwen 3.5 9B head; Flash-Next sources convert (PR 508). [source]
- A checkpoint with every trunk shard but no draft head inspects as autoregressive with MTP off, and Forge refuses to build a speculative artifact from it. [source]
- A model that ships as one weights file has no index, and Forge once converted it, wrote no `mtp.safetensors`, and stopped at calibration with `return_hidden requires an MTP-patched runtime`. [source]
- A head cannot be validated apart from the trunk it was trained against; a fine-tuned trunk can drift from the base head. [source]
- Draft quantization is contested (see the norm-convention dossier); the Flash-Next 8-bit recipe quantizes the head the same as the body. [source]
- What the README's "train the MTP adapter" does: gradient fine-tuning of an existing head against a changed trunk, or only calibration of wiring. No cached page gives data, steps, loss or hardware. [source]
- Whether any Mac-side route trains a head for a model with none (DeepSeek-style MTP head from scratch, or EAGLE-style speculator training) exists; the cached pages describe training only on CUDA hardware, if at all. [source]
- Whether grafting a base model's head onto a fine-tune keeps acceptance: MiMo's speed on MTPLX is listed as not yet measured. [source]
- The first MTPLX changelog entry for Forge reads "convert any Hugging Face repo to MLX (AWQ, compressed-tensors, NVFP4, BF16 sources), calibrate and train the MTP adapter, verify with quality gates that reject speed wins that degrade output, and publish with provenance". [source]
- The README says MTPLX does not support attaching a separately supplied MTP sidecar to an arbitrary MLX trunk, because matching fields, shapes or provenance labels cannot prove the head was trained against those exact trunk weights. [source]
- The README's "what MTPLX is not" says it is not an external-drafter system: the drafter is the target model's own MTP heads. [source]
- MTPLX PR 489 made `mtplx inspect` report a Qwen checkpoint with every trunk shard but no draft head (`nex-agi/Nex-N2.5-mini`) as runnable with MTP off, and Forge now requires MTP weight evidence because its runtime probe had treated `can_run` as MTP evidence and admitted AR-only trunks into a "speculative conversion that cannot create the missing trained head". [source]
- The README documents a Laguna model with no native MTP head that is run with `--no-mtp`, and an MTP launch is rejected before weights load instead of falling back during execution. [source]
- A changelog entry thanks a contributor for issue 218, which removed unsupported MTP-sidecar graft guidance and added bracketed corrections to historical release notes that documented commands that never worked; it was seeded by issue 215's phantom graft script. [source]
- Forge extracts MTP heads stored outside the `mtp.` prefix (PR 442), covering the GLM-4 MoE, GLM-5.3-Flash, DeepSeek-V3.2 and MiMo layouts, and binds MiMo's output head instead of leaving it random. [source]
- Xiaomi's MiMo-V2.6-Distill-Qwen-9B is a Qwen3.5-9B fine-tune whose MTPLX pack carries "the Qwen 3.5 9B MTP head", and its speed on MTPLX is listed as not yet measured. [source]
- A Forge calibration once switched a grafted Qwen 3.5 9B head to `pre_norm` on one prompt and six tokens in which every option accepted nothing; a switch now needs at least 4 prompts, 48 draft rounds per option and a lead of 0.25 accepted tokens per round, an inconclusive calibration is recorded in `mtplx_runtime.json`, and the probe samples 8 prompts and 4 windows (49 s on the Qwen 3.5 9B pack including model load). [source]
- In one MTPLX verification, contract calibration over 64 candidates on a broken artifact returned `best_agreement: 0.0` and `no_agreement_signal`, and the same calibration on a healthy artifact returned 1.0, which shows calibration selects among wiring candidates by measured agreement. [source]
- Issue 176's chain probe swept `hidden_variant` over post_norm, pre_norm and fc against history modes (recursive, target_forced and others) to localise a zero-acceptance head, which is a wiring search, not training. [source]
- A model small enough for a single `model.safetensors` has no index file, and Forge once converted it, wrote no `mtp.safetensors`, and stopped at calibration with `return_hidden requires an MTP-patched runtime` (issue 492) until Forge read the same file headers `inspect` reads. [source]
- Forge's Flash-Next Optimized-Quality recipe quantizes the main model and the draft head at 8 bits with group size 64, keeps structural weights in BF16 and the n-gram table at 4 bits with group size 32 (the fixed-width verifier rejects an 8-bit table). [source]
- mlx-lm PR 1468 attaches a separately published MTP drafter checkpoint such as `mlx-community/Qwen3.6-27B-MTP-bf16` to an already loaded target, borrowing the target's embedding table and output head and feeding it the target's hidden state, and lazily builds the head only when a checkpoint bundles `mtp.` weights. [source]
- The vllm-mlx conversion script described in mlx-lm issue 872 extracts BF16 MTP weights from the original Hugging Face model, selectively quantizes them to match the base model, and keeps norms and the `fc` projection in fp16, which is conversion of an existing head, not training. [source]
- No cached page describes training an MTP head, EAGLE speculator or other drafter on a Mac for a model that ships none. [source]
Corrections and disagreements
- README: "Forge ... convert to MLX, train the MTP adapter, verify". PR 489's changelog entry: a speculative conversion "cannot create the missing trained head". CONTRADICTS: mtplx-forge-and-mtp-preserving-mlx-conversion.md, whose Definition says Forge "keeps or trains an MTP head": the cached changelog shows Forge cannot train a head where none exists. The README's "train" may mean fitting or calibrating an existing head, but no cached page says so explicitly. [source]
Children
- No children recorded.