<!-- llms-explorer concept facts · https://llms-explorer.com/tree/mtplx-sustained-versus-turbo-presets-and-which-u/ · pack 2026-10-05 · ~924 tokens -->

# MTPLX Sustained versus Turbo presets and which use compiled verify

> Turbo is selected automatically for the quantized 27B and 9B flagship models; Sustained for the rest.

Parent: [Mac local LLMs: Speculative decoding and MTP](https://llms-explorer.com/tree/mac-local-llms-speculative-decoding-and-mtp/) · 2 facets · 19 facts · page: https://llms-explorer.com/tree/mtplx-sustained-versus-turbo-presets-and-which-u/

## Facts

- Turbo is selected automatically for the quantized 27B and 9B flagship models; Sustained for the rest. — source: `asserted`
- Compiled verify is a separate mechanism from the preset: Bonsai, which the README catalog does not list as Turbo, uses the compiled verifier. — source: `asserted`
- 2.0.0: the turbo profile defaults to compiled verify with a per-model gate (existing dossier). — source: `asserted`
- 2.12.0: Bonsai moves onto the compiled verifier after matching eager in 613 rounds. — source: `asserted`
- Compiled verify has no route for image-bearing context on Flash-Next; those turns fall back to eager (33-50 against 54-61 tok/s text-only), and 2.12.0 reports compiled with an image at 101.6 against 97.2 tok/s. — source: `asserted`
- Eager BF16 rounds are used for contexts past the compiled route's 32,768 tokens and for buffer-growth fallbacks. — source: `asserted`
- Whether the Sustained packs (Qwen 3.5 4B, MiMo 9B, Gemma 4) use compiled verify. No source states it. — source: `asserted`
- Bonsai's preset label is not in the README pack table. — source: `asserted`
- The README pack table lists Sustained for Qwen3.5-4B Optimized Speed and Quality (depth 3), MiMo-V2.6-Qwen-9B (depth 2) and Gemma4 Optimized Speed. — [source](https://github.com/youssofal/MTPLX)
- The README lists Turbo for Qwen3.5-9B (tuning on a 16 GB M4 Mac mini lands on depth 1), all Qwen3.8-27B packs at depth 3, and all three Flash-Next packs at depth 3. — [source](https://github.com/youssofal/MTPLX)
- Flash-Next Optimized Speed is listed as Turbo, depth 3, and the family accepts up to depth 5. — [source](https://github.com/youssofal/MTPLX)
- Gemma4 runs as an assistant pair, so its tuned control is the draft block size and not depth. — [source](https://github.com/youssofal/MTPLX)
- Bonsai now uses the compiled verifier and matched eager verification in all 613 rounds checked. — [source](https://mtplx.com/releases/2.12.0/)
- Sending an image to Flash-Next used to force eager verification for the whole conversation; 2.12.0 gives compiled verification at 101.6 against 97.2 tok/s. — [source](https://mtplx.com/releases/2.12.0/)
- Eager BF16 rounds apply to image requests, contexts past the compiled route's 32,768 tokens, buffer-growth fallbacks, warm-restore rounds and single eager passes inside a compiled request. — [source](https://mtplx.com/releases/2.12.0/)
- Sustained Max is Sustained with fans pinned at 100 percent, and a detached watchdog restores automatic fans if MTPLX dies. — [source](https://github.com/youssofal/MTPLX)
- A pack can set `recommended_generation_mode` to plain decoding while still shipping its draft head; no shipped pack uses it at 2.12.0. — [source](https://mtplx.com/releases/2.12.0/)
- A pack whose accuracy measurement is unpublished shows "Official MTPLX pack, qualification pending" and exits 0 from `mtplx inspect`; Bonsai 2 shows as verified. — [source](https://mtplx.com/releases/2.12.0/)

## Corrections and disagreements

- CONTRADICTS: mtplx-flash-decoding-verify-attention-kernel-sdp.md line "the README names compiled verify only for the Turbo preset" as implying Sustained packs never compile: 2.12.0 says Bonsai uses compiled verify, so compile status is per pack and not purely per preset. — source: `asserted`
