Upstreaming Prism ternary types 142/143 into llama.cpp and Ollama
Parent: Mac local LLMs: Quantization formats and methods · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Prism's Bonsai-demo README says Bonsai 2 support is being upstreamed in smaller PRs that target the official `Q2_0` format and that Bonsai 2 still requires the PrismML fork.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Prism's Bonsai-demo README says Bonsai 2 support is being upstreamed in smaller PRs that target the official `Q2_0` format and that Bonsai 2 still requires the PrismML fork. [source]
- The README status table, checked 2026-09-25, lists llama.cpp PRs 27779 and 29094 and 29095 as merged, 29096 and 29243 as open, and 29100 and 29101 as draft. [source]
- The README says PQ2_0 and PTQ1_0 "remain fork-specific packings" and that the upstream effort "does not promise support for those types". [source]
- llama.cpp PR 29094 makes the Metal FWHT kernel accept an F16 source and gates `supports_op` to the four kernel sizes. [source]
- llama.cpp PR 29096 does the same for CUDA, with `test-backend-ops -o MUL_MAT_HADAMARD` 15/15. [source]
- Prism's formats page says Ternary Bonsai 2 needs an activation-side Walsh-Hadamard transform that is "not upstream yet", so every backend needs the fork binaries `prism-b10658` or newer. [source]
- The formats page says Ollama and other tools that bundle stock llama.cpp cannot run Ternary Bonsai 2. [source]
- The formats page says the MLX package for Ternary Bonsai 2 uses a 2-bit group-128 packing analogous to PQ2_0 and runs on stock MLX. [source]
- A llama.cpp member commented on issue 29058 on 2026-09-19 with only "See also #22019". [source]
- An issue 29058 commenter reports master's `qwen35.cpp` MTP path reads `token_embd` with a raw `ggml_get_rows`, and a folded model with `--spec-type draft-mtp` fails with `Hadamard-latent table 'token_embd.weight' is read without the inverse transform`. [source]
- The same commenter says the Prism fork fixed this in commit 422590f5 on 2026-09-21. [source]
- The same commenter measured a standalone MTP sidecar at 5.16 t/s against 16.6 t/s without speculation, with rho 0.43 versus about 0.06 for a shared table. [source]
- The same commenter measured MTP acceptance on the 27B pack at 65.8%, 52.2%, 67.9%, 82.3% and 84.1% for 8k, 32k, 64k, 128k and 191k tokens of context. [source]
- The same commenter says folding the terminal norm gain `gf / rms(gf)` into `eh_proj` raised 0.8B acceptance from 10.7% to 28.4% and 27B from 35.6% to 40.5%. [source]
- The same commenter says an m=3 file can be served as m=2 without requantizing by dropping a plane. [source]
- The Prism formats page lists PTQ1_0 at 1.76 bpw and 5.93 GB and PQ2_0 at 2.16 bpw and 7.25 GB, which differs from the Hugging Face card's 1.75 bpw, 5.95 GB and 2.13 bpw, 7.21 GB. [source]
Children
- No children recorded.