Silent wrong-output loading of Hadamard-rotated F16 GGUF in stock llama.cpp
Parent: Mac local LLMs: Quantization formats and methods · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Prism's formats page says Do not run Ternary Bonsai 2 on stock llama.cpp because a `Q2_0` file "loads without a warning and outputs gibberish", which is why that band is kept in a separate repo.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Prism's formats page says Do not run Ternary Bonsai 2 on stock llama.cpp because a `Q2_0` file "loads without a warning and outputs gibberish", which is why that band is kept in a separate repo. [source]
- The Bonsai-demo README says the development `Q2_0` checkpoint lives in `Ternary-Bonsai-2-27B-gguf-dev` and still requires the Prism fork. [source]
- The Prism llama.cpp page says a `Q2_0` file loads silently on stock builds and outputs gibberish, and that PTQ1_0 and PQ2_0 files are refused. [source]
- The Hugging Face card says stock llama.cpp loads `Q2_0` without any warning and produces garbage because it has no Hadamard activation runtime. [source]
- Prism's formats page marks legacy `*-Q2_0.gguf` (no `g64`) as deprecated group-128 files under the id that now belongs to group-64, readable only by old `prism-v5` releases, with newer binaries refusing them. [source]
- The upstream FWHT PRs add the backend op with F16 input (CPU PR 27779, Metal PR 29094) and do not add a model-side Hadamard path. [source]
- On the Prism fork a folded model run with `--spec-type draft-mtp` is refused with `Hadamard-latent table 'token_embd.weight' is read without the inverse transform` instead of running silently. [source]
Children
- No children recorded.