KV-cache quantization degrading tool-call accuracy
Parent: Mac local LLMs: KV cache sizing and quantization · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Errors in cached keys perturb attention scores on every later token, so a tool call's exact argument strings (paths, ids copied from earlier turns) are the content most exposed to retrieval error; free prose tolerates it better. This is an inference, not a measured result.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Errors in cached keys perturb attention scores on every later token, so a tool call's exact argument strings (paths, ids copied from earlier turns) are the content most exposed to retrieval error; free prose tolerates it better. This is an inference, not a measured result. [source]
- Hybrid models with few attention layers (Qwen3.5/3.6: about 16 of 64 layers keep a KV cache) have little KV to compress, so quantization effects are small there and the saving is small too. [source]
- 2026-03: TurboQuant discussion (llama.cpp #20969) drove sub-4-bit KV proposals, mostly tested with perplexity and needle-in-a-haystack, not with tool calls. [source]
- 2026: guides for agentic coding recommend q8_0 K and V (OpenCode on 16 GB VRAM) while warning that K-cache quantization can produce wrong diffs. [source]
- Users running q4_0 K and V on small VRAM report the model "going off the rails" more once speculative decoding (MTP) is enabled, and trade speed for fewer errors; no isolation of KV quantization from MTP was done. [source]
- A smoke battery on a TurboQuant CUDA fork (RTX 5090, Qwen3.5-27B hybrid) passed a BFCL-style irrelevance check (model correctly made no tool call), 5/5 needle depths, RULER variable tracking, and byte-identical MTP output; that is one model, one sample per probe. [source]
- Stock TurboQuant implementations returned 0% needle retrieval on an M1 Pro until math and Metal fixes landed, so unvalidated KV code can fail silently well before tool calls are tested. [source]
- Safe or not: Rigel's guide calls q8_0 "practically lossless for most tasks" and uses it as its agentic-coding default, yet lists "tool-call errors" and incorrect diffs as a risk of K-cache quantization. HN commenters run q4_0 and q8_0 KV on Macs and call the result usable but "dumb"; they cannot separate it from weight quantization (a commenter points out the quant levels were unreported). [source]
- Which cache matters: the Rigel article says K and V can be quantized differently for different use cases; llama.cpp on Metal requires K and V types to match for fused attention (existing dossier), so the K/V split advice may not run fused on a Mac. [source]
- A controlled BFCL or agent-loop run varying only `-ctk/-ctv` (f16, q8_0, q4_0) on a fixed weight quant, on Metal. Not found anywhere. [source]
- Does attention rotation (llama.cpp PR 21038) shift tool-call accuracy at q4_0, as it shifts perplexity? [source]
- A Qwen 3.6-class hybrid model keeps KV cache in only about 16 of 64 layers, so TurboQuant added roughly 1 to 3% overhead and had little cache to compress. [source]
- A TurboQuant CUDA fork test on Qwen3.5-27B reported a BFCL-style irrelevance probe passing (plain text, empty tool_calls) alongside NIAH 5/5 and byte-identical MTP output. [source]
- The same discussion reports MTP tool-call generation at 102.03 tok/s versus 47.84 tok/s baseline on that fork, with 60.9% draft acceptance on tool calls. [source]
- An M1 Pro 16 GB TurboQuant evaluation found stock implementations scored 0% needle-in-haystack at every context length tried until QJL and Metal kernel fixes were applied. [source]
- A Hacker News commenter runs a 12B model with `-ctv q4_0 -ctk q4_0` and says MTP speed versus "less dumb" output is a trade-off he has not settled. [source]
- Another HN commenter points out that comparing local-model experiences is "apples to oranges" without knowing the weight and KV-cache quantization levels used. [source]
- An HN commenter on 8 GB VRAM driving Home Assistant by voice found a quantized Gemma 4 the best pick for many tool calls, with Qwen3.5 9B close behind. [source]
- The Rigel Computer guide says switching OpenCode's llama.cpp cache from f16 to q8_0 was its agentic-coding breakthrough, using `cache_type_k` and `cache_type_v` both q8_0 with qwen2.5-coder-14b Q5_K_M at 32k context in 16 GB VRAM. [source]
- The same guide warns that K-cache quantization can cause incorrect diffs and tool-call errors even in coding agents, without a measurement. [source]
- No source found compares tool-call accuracy at f16 versus q8_0 versus q4_0 KV cache with the weights held fixed. [source]
Children
- No children recorded.