<!-- llms-explorer concept facts · https://llms-explorer.com/tree/llama-cpp-attention-rotation-hadamard-for-quanti/ · pack 2026-10-05 · ~651 tokens -->

# llama.cpp attention rotation Hadamard for quantized KV cache

> Maintainer AIME25 test (gpt-oss-20b, low reasoning, 8 runs, 240 attempts): F16 KV 37.9%; Q8_0 31.7% without rotation vs 37.1% with; Q5_1 30.8 vs 32.5; Q5_0 25.4 vs 32.5; Q4_1 18.3 vs 28.3; Q4_0 2.0 vs 21.7.

Parent: [Mac local LLMs: KV cache sizing and quantization](https://llms-explorer.com/tree/mac-local-llms-kv-cache-sizing-and-quantization/) · 2 facets · 7 facts · page: https://llms-explorer.com/tree/llama-cpp-attention-rotation-hadamard-for-quanti/

## Facts

- Maintainer AIME25 test (gpt-oss-20b, low reasoning, 8 runs, 240 attempts): F16 KV 37.9%; Q8_0 31.7% without rotation vs 37.1% with; Q5_1 30.8 vs 32.5; Q5_0 25.4 vs 32.5; Q4_1 18.3 vs 28.3; Q4_0 2.0 vs 21.7. — [source](https://github.com/ggml-org/llama.cpp/pull/21038)
- In that test the Q4_0 no-rotation run solved essentially no problems while still producing coherent reasoning, i.e. the failure was wrong answers rather than garbage output; the maintainer saw non-negligible benefit even for Q8_0 and was still deciding whether to disable rotation there. — [source](https://github.com/ggml-org/llama.cpp/pull/21038)
- A downstream port applies rotation only when the K (resp. V) cache type is quantized and the head size is a multiple of 64, ports the ggml-org#21586 check that the quantized block size divides the head size when flash attention is on, and keeps rotation off for MLA models. — [source](https://github.com/ggml-org/llama.cpp/pull/21038)
- An Ollama branch adds `OLLAMA_KV_ROTATE=1` (default off, byte-identical behaviour when unset) rotating Q/K/V with a fixed normalized 64x64 Sylvester Hadamard block-diagonally along dim 0 when head dim is a multiple of 64. — [source](https://github.com/ggml-org/llama.cpp/pull/21038)
- A later Ollama-branch commit aligns K rotation to 64x64 instead of the largest power-of-two divisor (e.g. 128x128 for Llama-3), citing V's empirical result in #21038 that smaller matrices are better. — [source](https://github.com/ggml-org/llama.cpp/pull/21038)
- The same port notes its K-shift graph rotates back before RoPE and forward again afterwards, so context shifting stays consistent with a rotated cache. — [source](https://github.com/ggml-org/llama.cpp/pull/21038)

## Corrections and disagreements

- CONTRADICTS kv-cache-quantization-tradeoffs-on-apple-gpus.md (q4_0 near-lossless after rotation, +2.5% PPL): on a reasoning benchmark rotated Q4_0 still loses 16 points of AIME25 (21.7% vs 37.9% F16); perplexity understates agentic or reasoning damage. — [source](https://github.com/ggml-org/llama.cpp/pull/21038)
