Mixed-precision KV-cache quantization per-layer sensitivity
Parent: Mac local LLMs: KV cache sizing and quantization · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
KVTuner (ICML 2025 poster) searches layer-wise KV precision pairs offline and reports nearly lossless 3.25-bit mixed precision KV for Llama-3.1-8B-Instruct and 4.0-bit for the sensitive Qwen2.5-7B-Instruct on mathematical reasoning.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- KVTuner (ICML 2025 poster) searches layer-wise KV precision pairs offline and reports nearly lossless 3.25-bit mixed precision KV for Llama-3.1-8B-Instruct and 4.0-bit for the sensitive Qwen2.5-7B-Instruct on mathematical reasoning. [source]
- KVTuner reports up to 21.25% higher maximum inference throughput than KIVI-KV8 over various context lengths. [source]
- KVTuner's search uses intra-layer KV precision-pair pruning and inter-layer clustering to shrink the space, and the resulting per-layer configuration is applied offline-searched at inference with no online fine-grained decision overhead. [source]
- KVTuner finds a 5-bit K8V2 pair matches or beats a 6-bit K4V8 pair on generation quality while using 12.5% less memory; most models degrade only at int2 keys, but Qwen2.5-7B and Qwen2.5-Math-7B are sensitive even at int4 keys. [source]
- KVTuner finds per-channel asymmetric quantization beats per-token asymmetric for keys because keys have channel-wise outliers, and the Pareto-optimal precision pairs differ between the two modes. [source]
- KVTuner's account of why keys matter more: error accumulates across layers and steps and shifts per-layer attention distributions; when fine-grained token or page-level quantization is infeasible, raising key precision in sensitive layers (shrinking q·ΔK) is the recommended fix. [source]
- vLLM documents per-layer exclusion (`--kv-cache-dtype-skip-layers` by layer index or `sliding_window`) and notes per-attention-head FP8 scales exist only with the Flash Attention backend and llm-compressor calibration. [source]
- Hannecke's TurboQuant-on-macOS survey says its per-layer sensitivity analysis shows not all layers benefit equally and that all 10 full-attention layers of Qwen3.5 need individual bit allocation depending on weight quantization, while no fused Metal dequant kernel inside flash attention existed for that path. [source]
- Inference: on Macs the practical per-layer lever today is coarse (exclude SWA layers, protect boundary layers), not a searched per-layer map, because llama.cpp and mlx-lm expose one K type and one V type globally. [source]
Children
- No children recorded.