<!-- llms-explorer concept facts · https://llms-explorer.com/tree/mixed-precision-kv-cache-quantization-per-layer/ · pack 2026-10-05 · ~788 tokens -->

# Mixed-precision KV-cache quantization per-layer sensitivity

> KVTuner (ICML 2025 poster) searches layer-wise KV precision pairs offline and reports nearly lossless 3.25-bit mixed precision KV for Llama-3.1-8B-Instruct and 4.0-bit for the sensitive Qwen2.5-7B-Instruct on mathematical reasoning.

Parent: [Mac local LLMs: KV cache sizing and quantization](https://llms-explorer.com/tree/mac-local-llms-kv-cache-sizing-and-quantization/) · 1 facets · 9 facts · page: https://llms-explorer.com/tree/mixed-precision-kv-cache-quantization-per-layer/

## Facts

- KVTuner (ICML 2025 poster) searches layer-wise KV precision pairs offline and reports nearly lossless 3.25-bit mixed precision KV for Llama-3.1-8B-Instruct and 4.0-bit for the sensitive Qwen2.5-7B-Instruct on mathematical reasoning. — [source](https://openreview.net/forum?id=zDwipF6h06)
- KVTuner reports up to 21.25% higher maximum inference throughput than KIVI-KV8 over various context lengths. — [source](https://openreview.net/forum?id=zDwipF6h06)
- KVTuner's search uses intra-layer KV precision-pair pruning and inter-layer clustering to shrink the space, and the resulting per-layer configuration is applied offline-searched at inference with no online fine-grained decision overhead. — [source](https://arxiv.org/html/2502.04420v5)
- KVTuner finds a 5-bit K8V2 pair matches or beats a 6-bit K4V8 pair on generation quality while using 12.5% less memory; most models degrade only at int2 keys, but Qwen2.5-7B and Qwen2.5-Math-7B are sensitive even at int4 keys. — [source](https://arxiv.org/html/2502.04420v5)
- KVTuner finds per-channel asymmetric quantization beats per-token asymmetric for keys because keys have channel-wise outliers, and the Pareto-optimal precision pairs differ between the two modes. — [source](https://arxiv.org/html/2502.04420v5)
- KVTuner's account of why keys matter more: error accumulates across layers and steps and shifts per-layer attention distributions; when fine-grained token or page-level quantization is infeasible, raising key precision in sensitive layers (shrinking q·ΔK) is the recommended fix. — [source](https://arxiv.org/html/2502.04420v5)
- vLLM documents per-layer exclusion (`--kv-cache-dtype-skip-layers` by layer index or `sliding_window`) and notes per-attention-head FP8 scales exist only with the Flash Attention backend and llm-compressor calibration. — [source](https://docs.vllm.ai/en/latest/features/quantization/quantized_kvcache/)
- Hannecke's TurboQuant-on-macOS survey says its per-layer sensitivity analysis shows not all layers benefit equally and that all 10 full-attention layers of Qwen3.5 need individual bit allocation depending on weight quantization, while no fused Metal dequant kernel inside flash attention existed for that path. — [source](https://medium.com/@michael.hannecke/turboquant-on-apple-macos-five-integration-paths-for-local-kv-cache-compression-42e83959d414)
- Inference: on Macs the practical per-layer lever today is coarse (exclude SWA layers, protect boundary layers), not a searched per-layer map, because llama.cpp and mlx-lm expose one K type and one V type globally. — source: `asserted`
