MTPLX Sustained versus Turbo presets and which use compiled verify
Parent: Mac local LLMs: Speculative decoding and MTP · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Turbo is selected automatically for the quantized 27B and 9B flagship models; Sustained for the rest.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Turbo is selected automatically for the quantized 27B and 9B flagship models; Sustained for the rest. [source]
- Compiled verify is a separate mechanism from the preset: Bonsai, which the README catalog does not list as Turbo, uses the compiled verifier. [source]
- 2.0.0: the turbo profile defaults to compiled verify with a per-model gate (existing dossier). [source]
- 2.12.0: Bonsai moves onto the compiled verifier after matching eager in 613 rounds. [source]
- Compiled verify has no route for image-bearing context on Flash-Next; those turns fall back to eager (33-50 against 54-61 tok/s text-only), and 2.12.0 reports compiled with an image at 101.6 against 97.2 tok/s. [source]
- Eager BF16 rounds are used for contexts past the compiled route's 32,768 tokens and for buffer-growth fallbacks. [source]
- Whether the Sustained packs (Qwen 3.5 4B, MiMo 9B, Gemma 4) use compiled verify. No source states it. [source]
- Bonsai's preset label is not in the README pack table. [source]
- The README pack table lists Sustained for Qwen3.5-4B Optimized Speed and Quality (depth 3), MiMo-V2.6-Qwen-9B (depth 2) and Gemma4 Optimized Speed. [source]
- The README lists Turbo for Qwen3.5-9B (tuning on a 16 GB M4 Mac mini lands on depth 1), all Qwen3.8-27B packs at depth 3, and all three Flash-Next packs at depth 3. [source]
- Flash-Next Optimized Speed is listed as Turbo, depth 3, and the family accepts up to depth 5. [source]
- Gemma4 runs as an assistant pair, so its tuned control is the draft block size and not depth. [source]
- Bonsai now uses the compiled verifier and matched eager verification in all 613 rounds checked. [source]
- Sending an image to Flash-Next used to force eager verification for the whole conversation; 2.12.0 gives compiled verification at 101.6 against 97.2 tok/s. [source]
- Eager BF16 rounds apply to image requests, contexts past the compiled route's 32,768 tokens, buffer-growth fallbacks, warm-restore rounds and single eager passes inside a compiled request. [source]
- Sustained Max is Sustained with fans pinned at 100 percent, and a detached watchdog restores automatic fans if MTPLX dies. [source]
- A pack can set `recommended_generation_mode` to plain decoding while still shipping its draft head; no shipped pack uses it at 2.12.0. [source]
- A pack whose accuracy measurement is unpublished shows "Official MTPLX pack, qualification pending" and exits 0 from `mtplx inspect`; Bonsai 2 shows as verified. [source]
Corrections and disagreements
- CONTRADICTS: mtplx-flash-decoding-verify-attention-kernel-sdp.md line "the README names compiled verify only for the Turbo preset" as implying Sustained packs never compile: 2.12.0 says Bonsai uses compiled verify, so compile status is per pack and not purely per preset. [source]
Children
- No children recorded.