<!-- llms-explorer concept facts · https://llms-explorer.com/tree/prism-ml-ternary-hadamard-rotated-weights-kernel/ · pack 2026-10-05 · ~829 tokens -->

# Prism ML ternary Hadamard-rotated weights kernel on Metal

> Ternary Bonsai 2 27B weights are rotated blockwise by an orthogonal Hadamard transform (block 1024, fixed plus/minus 1 signs) before ternary assignment, and the runtime applies the matching transform to activations

Parent: [Mac local LLMs: MLX kernels, numerics and internals](https://llms-explorer.com/tree/mac-local-llms-mlx-kernels-numerics-and-internals/) · 1 facets · 14 facts · page: https://llms-explorer.com/tree/prism-ml-ternary-hadamard-rotated-weights-kernel/

## Facts

- Ternary Bonsai 2 27B weights are rotated blockwise by an orthogonal Hadamard transform (block 1024, fixed plus/minus 1 signs) before ternary assignment, and the runtime applies the matching transform to activations — [source](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-mlx-2bit)
- The rotation is folded into the stored weights offline and costs no extra bits or weight traffic — [source](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-mlx-2bit)
- The packed model declares its rotation as metadata so a runtime must apply the matching transform or refuse to load — [source](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-mlx-2bit)
- Ordinary MLX loaders skip the activation transform and the inverse embedding lookup and return wrong output rather than an error — [source](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-mlx-2bit)
- The weight format is ternary g128 with one FP16 scale per 128 weights, about 1.71 bits per weight, 1.72 for the whole language model — [source](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-mlx-2bit)
- MLX stores the ternary levels as 2-bit codes with scale s and bias -s, which costs 2.25 bits per weight against 2.13 for PQ2_0 — [source](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-mlx-2bit)
- The MLX pack is 8.60 GB: 7.67 GB language model plus a 0.92 GB unquantized FP16 vision tower that needs no rotation — [source](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-mlx-2bit)
- 26.2M parameters (0.0976 percent of the language model) stay in higher precision — [source](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-mlx-2bit)
- Two GGUF packings exist: PTQ1_0 at 5.95 GB and PQ2_0 at 7.21 GB, and neither is uniformly faster across GPUs — [source](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-mlx-2bit)
- The model card's M5 Pro Metal row (llama.cpp fork, PQ2_0) is 28.1 tok/s TG128 and 387 tok/s PP512 with decode streaming about 204 GB/s of weights and drawing 27.5 W on the GPU rail — [source](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-mlx-2bit)
- Earlier pre-rotation Apple numbers: M5 Max 47.0 TG128 and 765 PP512, M4 Pro 18.0 TG128 and 125 PP512, all at 7.2 GB — [source](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-mlx-2bit)
- Stock llama.cpp cannot run the Bonsai 2 GGUF packs; the PrismML-Eng/llama.cpp fork is required — [source](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-mlx-2bit)
- Prism ML states 1.76 effective bits, 5.9 GB, 98.2 percent aggregate retention (83.9 against 85.4 for Qwen3.8 27B) and 46.8 tok/s on M5 Max — [source](https://prismml.com/news/bonsai-2-27b)
- The announcement is dated 2026-09-17 and says the first Ternary Bonsai 27B retained about 95 percent — [source](https://prismml.com/news/bonsai-2-27b)
