<!-- llms-explorer concept facts · https://llms-explorer.com/tree/ternary-bonsai-2-27b-on-mtplx-with-grafted-qwen3/ · pack 2026-10-05 · ~796 tokens -->

# Ternary Bonsai 2 27B on MTPLX with grafted Qwen3.8 draft head

> Two kernels written for the pack: a one-step Prism Hadamard rotation (same bits as the four operations it replaces) and a ternary decode matmul.

Parent: [Mac local LLMs: Speculative decoding and MTP](https://llms-explorer.com/tree/mac-local-llms-speculative-decoding-and-mtp/) · 1 facets · 18 facts · page: https://llms-explorer.com/tree/ternary-bonsai-2-27b-on-mtplx-with-grafted-qwen3/

## Facts

- Two kernels written for the pack: a one-step Prism Hadamard rotation (same bits as the four operations it replaces) and a ternary decode matmul. — source: `asserted`
- On every load the kernels are tested on that Mac's GPU against stock ops; a failing kernel is disabled until restart and `/health` reports it. — source: `asserted`
- 2.12.0 added the pack (#515); the loader is named `prism_hadamard_qwen35`. — source: `asserted`
- The 16 GB planner used to refuse Bonsai because it could not also fund a warm RAM cache; it now admits it with an 8,192-token window, and a 16K prompt peaks at 12.11 GiB, over the 12.0 GiB budget. — source: `asserted`
- `xhigh` reasoning spent 577 s and 21,848 reasoning tokens on a coding task without answering, where `medium` finished. — source: `asserted`
- Kernel speed on M1 to M4 (only the M5 Max has measured speed). — source: `asserted`
- The Bonsai head accepts 75 percent of drafts, and depth 3 was 14 percent slower than depth 1 over a 3,000-token answer, so depth 1 is the default. — [source](https://mtplx.com/releases/2.12.0/)
- Bonsai uses the compiled verifier, which matched eager in all 613 rounds checked. — [source](https://mtplx.com/releases/2.12.0/)
- The pack is 8.85 GB: 8.60 GB of Prism ML's language model and vision tower unchanged, plus 0.24 GB of draft head. — [source](https://mtplx.com/releases/2.12.0/)
- Served id is `mtplx-bonsai-2-27b-optimized-speed`. — [source](https://mtplx.com/releases/2.12.0/)
- Both kernels run on every Mac from M1 to M5, and only the M5 Max has measured speed. — [source](https://mtplx.com/releases/2.12.0/)
- The kernels also make prompt reading 7 to 8 percent faster. — [source](https://mtplx.com/releases/2.12.0/)
- Reasoning efforts are `medium` (default) and `xhigh`; Prism ML states `low` is unsupported. — [source](https://mtplx.com/releases/2.12.0/)
- Memory classes: 16 GB uses 12.0 GiB budget and an 8,192-token window; 18 GB 13.5 GiB and 20,480 tokens (36,864 with 8-bit KV); 24 GB 18.0 GiB and 94,208 tokens (167,936 with 8-bit KV). — [source](https://mtplx.com/releases/2.12.0/)
- With the draft head under the 16 GB budget, a 7,006-token prompt and 1,024-token answer peaked at 11.54 GiB GPU and 12.79 GiB process memory with no swap growth. — [source](https://mtplx.com/releases/2.12.0/)
- The peak sits about 3.1 GiB above weights and KV cache, which is what the planner sets aside. — [source](https://mtplx.com/releases/2.12.0/)
- Bonsai 2's accuracy measurement is published, so `mtplx inspect` shows it as verified. — [source](https://mtplx.com/releases/2.12.0/)
- An app menu fix raised Bonsai decode from 45.5 and 46.9 to 48.5 tok/s with the menu open. — [source](https://mtplx.com/releases/2.12.0/)
