<!-- llms-explorer concept facts · https://llms-explorer.com/tree/kv-cache-quantization-degrading-tool-call-accura/ · pack 2026-10-05 · ~1368 tokens -->

# KV-cache quantization degrading tool-call accuracy

> Errors in cached keys perturb attention scores on every later token, so a tool call's exact argument strings (paths, ids copied from earlier turns) are the content most exposed to retrieval error; free prose tolerates it better. This is an inference, not a measured result.

Parent: [Mac local LLMs: KV cache sizing and quantization](https://llms-explorer.com/tree/mac-local-llms-kv-cache-sizing-and-quantization/) · 1 facets · 21 facts · page: https://llms-explorer.com/tree/kv-cache-quantization-degrading-tool-call-accura/

## Facts

- Errors in cached keys perturb attention scores on every later token, so a tool call's exact argument strings (paths, ids copied from earlier turns) are the content most exposed to retrieval error; free prose tolerates it better. This is an inference, not a measured result. — source: `asserted`
- Hybrid models with few attention layers (Qwen3.5/3.6: about 16 of 64 layers keep a KV cache) have little KV to compress, so quantization effects are small there and the saving is small too. — source: `asserted`
- 2026-03: TurboQuant discussion (llama.cpp #20969) drove sub-4-bit KV proposals, mostly tested with perplexity and needle-in-a-haystack, not with tool calls. — source: `asserted`
- 2026: guides for agentic coding recommend q8_0 K and V (OpenCode on 16 GB VRAM) while warning that K-cache quantization can produce wrong diffs. — source: `asserted`
- Users running q4_0 K and V on small VRAM report the model "going off the rails" more once speculative decoding (MTP) is enabled, and trade speed for fewer errors; no isolation of KV quantization from MTP was done. — source: `asserted`
- A smoke battery on a TurboQuant CUDA fork (RTX 5090, Qwen3.5-27B hybrid) passed a BFCL-style irrelevance check (model correctly made no tool call), 5/5 needle depths, RULER variable tracking, and byte-identical MTP output; that is one model, one sample per probe. — source: `asserted`
- Stock TurboQuant implementations returned 0% needle retrieval on an M1 Pro until math and Metal fixes landed, so unvalidated KV code can fail silently well before tool calls are tested. — source: `asserted`
- Safe or not: Rigel's guide calls q8_0 "practically lossless for most tasks" and uses it as its agentic-coding default, yet lists "tool-call errors" and incorrect diffs as a risk of K-cache quantization. HN commenters run q4_0 and q8_0 KV on Macs and call the result usable but "dumb"; they cannot separate it from weight quantization (a commenter points out the quant levels were unreported). — source: `asserted`
- Which cache matters: the Rigel article says K and V can be quantized differently for different use cases; llama.cpp on Metal requires K and V types to match for fused attention (existing dossier), so the K/V split advice may not run fused on a Mac. — source: `asserted`
- A controlled BFCL or agent-loop run varying only `-ctk/-ctv` (f16, q8_0, q4_0) on a fixed weight quant, on Metal. Not found anywhere. — source: `asserted`
- Does attention rotation (llama.cpp PR 21038) shift tool-call accuracy at q4_0, as it shifts perplexity? — source: `asserted`
- A Qwen 3.6-class hybrid model keeps KV cache in only about 16 of 64 layers, so TurboQuant added roughly 1 to 3% overhead and had little cache to compress. — [source](https://github.com/ggml-org/llama.cpp/discussions/20969)
- A TurboQuant CUDA fork test on Qwen3.5-27B reported a BFCL-style irrelevance probe passing (plain text, empty tool_calls) alongside NIAH 5/5 and byte-identical MTP output. — [source](https://github.com/ggml-org/llama.cpp/discussions/20969)
- The same discussion reports MTP tool-call generation at 102.03 tok/s versus 47.84 tok/s baseline on that fork, with 60.9% draft acceptance on tool calls. — [source](https://github.com/ggml-org/llama.cpp/discussions/20969)
- An M1 Pro 16 GB TurboQuant evaluation found stock implementations scored 0% needle-in-haystack at every context length tried until QJL and Metal kernel fixes were applied. — [source](https://github.com/ggml-org/llama.cpp/discussions/20969)
- A Hacker News commenter runs a 12B model with `-ctv q4_0 -ctk q4_0` and says MTP speed versus "less dumb" output is a trade-off he has not settled. — [source](https://news.ycombinator.com/item?id=49402232)
- Another HN commenter points out that comparing local-model experiences is "apples to oranges" without knowing the weight and KV-cache quantization levels used. — [source](https://news.ycombinator.com/item?id=49402232)
- An HN commenter on 8 GB VRAM driving Home Assistant by voice found a quantized Gemma 4 the best pick for many tool calls, with Qwen3.5 9B close behind. — [source](https://news.ycombinator.com/item?id=49402232)
- The Rigel Computer guide says switching OpenCode's llama.cpp cache from f16 to q8_0 was its agentic-coding breakthrough, using `cache_type_k` and `cache_type_v` both q8_0 with qwen2.5-coder-14b Q5_K_M at 32k context in 16 GB VRAM. — [source](https://medium.com/rigel-computer-com/optimize-your-gpu-kv-cache-for-llama-cpp-opencode-co-13b6bc74f5ec)
- The same guide warns that K-cache quantization can cause incorrect diffs and tool-call errors even in coding agents, without a measurement. — [source](https://medium.com/rigel-computer-com/optimize-your-gpu-kv-cache-for-llama-cpp-opencode-co-13b6bc74f5ec)
- No source found compares tool-call accuracy at f16 versus q8_0 versus q4_0 KV cache with the weights held fixed. — source: `asserted`
