llama.cpp attention rotation Hadamard for quantized KV cache
Parent: Mac local LLMs: KV cache sizing and quantization · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Maintainer AIME25 test (gpt-oss-20b, low reasoning, 8 runs, 240 attempts): F16 KV 37.9%; Q8_0 31.7% without rotation vs 37.1% with; Q5_1 30.8 vs 32.5; Q5_0 25.4 vs 32.5; Q4_1 18.3 vs 28.3; Q4_0 2.0 vs 21.7.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Maintainer AIME25 test (gpt-oss-20b, low reasoning, 8 runs, 240 attempts): F16 KV 37.9%; Q8_0 31.7% without rotation vs 37.1% with; Q5_1 30.8 vs 32.5; Q5_0 25.4 vs 32.5; Q4_1 18.3 vs 28.3; Q4_0 2.0 vs 21.7. [source]
- In that test the Q4_0 no-rotation run solved essentially no problems while still producing coherent reasoning, i.e. the failure was wrong answers rather than garbage output; the maintainer saw non-negligible benefit even for Q8_0 and was still deciding whether to disable rotation there. [source]
- A downstream port applies rotation only when the K (resp. V) cache type is quantized and the head size is a multiple of 64, ports the ggml-org#21586 check that the quantized block size divides the head size when flash attention is on, and keeps rotation off for MLA models. [source]
- An Ollama branch adds `OLLAMA_KV_ROTATE=1` (default off, byte-identical behaviour when unset) rotating Q/K/V with a fixed normalized 64x64 Sylvester Hadamard block-diagonally along dim 0 when head dim is a multiple of 64. [source]
- A later Ollama-branch commit aligns K rotation to 64x64 instead of the largest power-of-two divisor (e.g. 128x128 for Llama-3), citing V's empirical result in #21038 that smaller matrices are better. [source]
- The same port notes its K-shift graph rotates back before RoPE and forward again afterwards, so context shifting stays consistent with a rotated cache. [source]
Corrections and disagreements
- CONTRADICTS kv-cache-quantization-tradeoffs-on-apple-gpus.md (q4_0 near-lossless after rotation, +2.5% PPL): on a reasoning benchmark rotated Q4_0 still loses 16 points of AIME25 (21.7% vs 37.9% F16); perplexity understates agentic or reasoning damage. [source]
Children
- No children recorded.