MTPLX Forge and MTP-preserving MLX conversion
Parent: Mac local LLMs: Speculative decoding and MTP · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
2.9.2 (25 Aug 2026): Forge correctness fix for the MTP norm convention.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- 2.9.2 (25 Aug 2026): Forge correctness fix for the MTP norm convention. [source]
- 2.12.0 (Sep 2026): Forge converts Flash-Next source models (issue 508). [source]
- The home page describes the flow as "paste a Hugging Face link". [source]
- A wrong norm convention collapsed draft acceptance to 0-2% on some packs until 2.9.2. [source]
- A trunk shifted twice (norm offset applied twice) is now refused by the runtime with a clear error. [source]
- A Forge verdict can be that MTP does not help: the tool reports "Depth 1 is fastest" style results from measurement rather than assuming a gain. [source]
- None found. All Forge claims are first-party (MTPLX README and release notes); no third party describes using Forge. [source]
- What data and how many steps Forge uses to train the MTP adapter is not in the cached pages. [source]
- Whether Forge can preserve an existing in-checkpoint MTP head losslessly, as against training a new adapter, is not stated. [source]
- Third-party reproduction of Forge speedups is missing. [source]
- MTPLX Forge turns a Hugging Face repo into an MTPLX-ready MTP model by converting to MLX, training the MTP adapter, verifying the result is faster and still exact, and optionally publishing to the Hub. [source]
- `mtplx forge --help` lists probe, build, publish and verify subcommands, and Forge is also a screen in the Mac app. [source]
- Pulls and Forge builds always write to the primary model root; extra `model_dirs` are discovery-only, and CLI updates or removals in an extra root need `--cache-dir` naming that root. [source]
- MTPLX 2.9.2 (25 Aug 2026) decides the MTP norm convention once per tensor set, which rescued packs whose draft acceptance had collapsed to 0-2%. [source]
- Since 2.9.2 the MTPLX runtime refuses a double-shifted trunk with a clear error instead of running it. [source]
- MTPLX 2.12.0 lets Forge convert Flash-Next source models (issue 508, contributed by bpmforge). [source]
- The MTPLX home page describes Forge as "paste a Hugging Face link; Forge converts it to MLX and measures the speedup on your Mac." [source]
- The official Youssofal catalog includes FP16 builds of Qwen 3.8 27B for M1 and M2, a Qwen 3.6 35B MoE balance build, and Gemma 4 packs. [source]
- A merge commit in the MTPLX repo lists `mtplx/commands/forge.py`, `compressed_tensors.py`, `gdn_capture.py` and `mtp_patch.py` as the conflicting modules of the 2.9.2 staging merge. [source]
- The MTPLX repo's CI fails any push whose history contains AI authorship or co-author attribution. [source]
Children
- No children recorded.