Ternary Bonsai 2 27B on MTPLX with grafted Qwen3.8 draft head
Parent: Mac local LLMs: Speculative decoding and MTP · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Two kernels written for the pack: a one-step Prism Hadamard rotation (same bits as the four operations it replaces) and a ternary decode matmul.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Two kernels written for the pack: a one-step Prism Hadamard rotation (same bits as the four operations it replaces) and a ternary decode matmul. [source]
- On every load the kernels are tested on that Mac's GPU against stock ops; a failing kernel is disabled until restart and `/health` reports it. [source]
- 2.12.0 added the pack (#515); the loader is named `prism_hadamard_qwen35`. [source]
- The 16 GB planner used to refuse Bonsai because it could not also fund a warm RAM cache; it now admits it with an 8,192-token window, and a 16K prompt peaks at 12.11 GiB, over the 12.0 GiB budget. [source]
- `xhigh` reasoning spent 577 s and 21,848 reasoning tokens on a coding task without answering, where `medium` finished. [source]
- Kernel speed on M1 to M4 (only the M5 Max has measured speed). [source]
- The Bonsai head accepts 75 percent of drafts, and depth 3 was 14 percent slower than depth 1 over a 3,000-token answer, so depth 1 is the default. [source]
- Bonsai uses the compiled verifier, which matched eager in all 613 rounds checked. [source]
- The pack is 8.85 GB: 8.60 GB of Prism ML's language model and vision tower unchanged, plus 0.24 GB of draft head. [source]
- Served id is `mtplx-bonsai-2-27b-optimized-speed`. [source]
- Both kernels run on every Mac from M1 to M5, and only the M5 Max has measured speed. [source]
- The kernels also make prompt reading 7 to 8 percent faster. [source]
- Reasoning efforts are `medium` (default) and `xhigh`; Prism ML states `low` is unsupported. [source]
- Memory classes: 16 GB uses 12.0 GiB budget and an 8,192-token window; 18 GB 13.5 GiB and 20,480 tokens (36,864 with 8-bit KV); 24 GB 18.0 GiB and 94,208 tokens (167,936 with 8-bit KV). [source]
- With the draft head under the 16 GB budget, a 7,006-token prompt and 1,024-token answer peaked at 11.54 GiB GPU and 12.79 GiB process memory with no swap growth. [source]
- The peak sits about 3.1 GiB above weights and KV cache, which is what the planner sets aside. [source]
- Bonsai 2's accuracy measurement is published, so `mtplx inspect` shows it as verified. [source]
- An app menu fix raised Bonsai decode from 45.5 and 46.9 to 48.5 tok/s with the menu open. [source]
Children
- No children recorded.