Prism ML ternary Hadamard-rotated weights kernel on Metal
Parent: Mac local LLMs: MLX kernels, numerics and internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Ternary Bonsai 2 27B weights are rotated blockwise by an orthogonal Hadamard transform (block 1024, fixed plus/minus 1 signs) before ternary assignment, and the runtime applies the matching transform to activations
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Ternary Bonsai 2 27B weights are rotated blockwise by an orthogonal Hadamard transform (block 1024, fixed plus/minus 1 signs) before ternary assignment, and the runtime applies the matching transform to activations [source]
- The rotation is folded into the stored weights offline and costs no extra bits or weight traffic [source]
- The packed model declares its rotation as metadata so a runtime must apply the matching transform or refuse to load [source]
- Ordinary MLX loaders skip the activation transform and the inverse embedding lookup and return wrong output rather than an error [source]
- The weight format is ternary g128 with one FP16 scale per 128 weights, about 1.71 bits per weight, 1.72 for the whole language model [source]
- MLX stores the ternary levels as 2-bit codes with scale s and bias -s, which costs 2.25 bits per weight against 2.13 for PQ2_0 [source]
- The MLX pack is 8.60 GB: 7.67 GB language model plus a 0.92 GB unquantized FP16 vision tower that needs no rotation [source]
- 26.2M parameters (0.0976 percent of the language model) stay in higher precision [source]
- Two GGUF packings exist: PTQ1_0 at 5.95 GB and PQ2_0 at 7.21 GB, and neither is uniformly faster across GPUs [source]
- The model card's M5 Pro Metal row (llama.cpp fork, PQ2_0) is 28.1 tok/s TG128 and 387 tok/s PP512 with decode streaming about 204 GB/s of weights and drawing 27.5 W on the GPU rail [source]
- Earlier pre-rotation Apple numbers: M5 Max 47.0 TG128 and 765 PP512, M4 Pro 18.0 TG128 and 125 PP512, all at 7.2 GB [source]
- Stock llama.cpp cannot run the Bonsai 2 GGUF packs; the PrismML-Eng/llama.cpp fork is required [source]
- Prism ML states 1.76 effective bits, 5.9 GB, 98.2 percent aggregate retention (83.9 against 85.4 for Qwen3.8 27B) and 46.8 tok/s on M5 Max [source]
- The announcement is dated 2026-09-17 and says the first Ternary Bonsai 27B retained about 95 percent [source]
Children
- No children recorded.