<!-- llms-explorer concept facts · https://llms-explorer.com/tree/silent-wrong-output-loading-of-hadamard-rotated/ · pack 2026-10-05 · ~526 tokens -->

# Silent wrong-output loading of Hadamard-rotated F16 GGUF in stock llama.cpp

> Prism's formats page says Do not run Ternary Bonsai 2 on stock llama.cpp because a `Q2_0` file "loads without a warning and outputs gibberish", which is why that band is kept in a separate repo.

Parent: [Mac local LLMs: Quantization formats and methods](https://llms-explorer.com/tree/mac-local-llms-quantization-formats-and-methods/) · 1 facets · 7 facts · page: https://llms-explorer.com/tree/silent-wrong-output-loading-of-hadamard-rotated/

## Facts

- Prism's formats page says Do not run Ternary Bonsai 2 on stock llama.cpp because a `Q2_0` file "loads without a warning and outputs gibberish", which is why that band is kept in a separate repo. — [source](https://docs.prismml.com/download/formats)
- The Bonsai-demo README says the development `Q2_0` checkpoint lives in `Ternary-Bonsai-2-27B-gguf-dev` and still requires the Prism fork. — [source](https://github.com/PrismML-Eng/Bonsai-demo)
- The Prism llama.cpp page says a `Q2_0` file loads silently on stock builds and outputs gibberish, and that PTQ1_0 and PQ2_0 files are refused. — [source](https://docs.prismml.com/run/llamacpp)
- The Hugging Face card says stock llama.cpp loads `Q2_0` without any warning and produces garbage because it has no Hadamard activation runtime. — [source](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf)
- Prism's formats page marks legacy `*-Q2_0.gguf` (no `g64`) as deprecated group-128 files under the id that now belongs to group-64, readable only by old `prism-v5` releases, with newer binaries refusing them. — [source](https://docs.prismml.com/download/formats)
- The upstream FWHT PRs add the backend op with F16 input (CPU PR 27779, Metal PR 29094) and do not add a model-side Hadamard path. — [source](https://github.com/PrismML-Eng/Bonsai-demo)
- On the Prism fork a folded model run with `--spec-type draft-mtp` is refused with `Hadamard-latent table 'token_embd.weight' is read without the inverse transform` instead of running silently. — [source](https://github.com/ggml-org/llama.cpp/issues/29058)
