<!-- llms-explorer concept facts · https://llms-explorer.com/tree/triaxialkv-per-role-mixed-precision-kv-quantizat/ · pack 2026-10-05 · ~672 tokens -->

# TriAxialKV per-role mixed-precision KV quantization for agentic prefills

> TriAxialKV tags tokens on three axes with 4 temporal values, 2 modal values and 7 semantic values, and combinations that never occur drop out of the optimization.

Parent: [Mac local LLMs: KV cache sizing and quantization](https://llms-explorer.com/tree/mac-local-llms-kv-cache-sizing-and-quantization/) · 1 facets · 11 facts · page: https://llms-explorer.com/tree/triaxialkv-per-role-mixed-precision-kv-quantizat/

## Facts

- TriAxialKV tags tokens on three axes with 4 temporal values, 2 modal values and 7 semantic values, and combinations that never occur drop out of the optimization. — [source](https://arxiv.org/html/2605.17170v1)
- The tagger runs on the CPU at scheduling time from chat-template special tokens, needing no model inference. — [source](https://arxiv.org/html/2605.17170v1)
- Calibration uses 5% of a dataset and hooks on four to six layers spread across depth, capturing prefill Q, K and V only. — [source](https://arxiv.org/html/2605.17170v1)
- Per-tag distortion D_k(b) is attention-output MSE, aggregated as max over heads, mean over requests, sum over layers; the paper argues output MSE is not monotonic in raw KV error. — [source](https://arxiv.org/html/2605.17170v1)
- The allocator solves INT2 or INT4 per tag under an average-bit budget, by exhaustive search for 22 or fewer tags and a greedy per-bit-gain rule otherwise. — [source](https://arxiv.org/html/2605.17170v1)
- The quantizer is asymmetric groupwise with group size 32, INT4 per-token for K and V, INT2 per-channel for keys and per-token for values. — [source](https://arxiv.org/html/2605.17170v1)
- Leftover INT2-key tokens that do not fill a group of 32 are routed through the INT4 per-token path rather than a full-precision residual buffer of 128 tokens per layer. — [source](https://arxiv.org/html/2605.17170v1)
- INT2 and INT4 page pools share one virtual address space split at an offset derived from the calibrated average bitwidth. — [source](https://arxiv.org/html/2605.17170v1)
- The decode path partitions each request's page table so INT2 pointers precede INT4, giving bitwidth-homogeneous flash-decoding splits merged by online log-sum-exp. — [source](https://arxiv.org/html/2605.17170v1)
- Reported throughput is 1.26x BF16 on Qwen3-VL-8B (B200), 1.32x on Qwen3-VL-32B (B200) and 1.52x on Qwen3-VL-32B (H100), with 3.4 to 4.0x more concurrent in-flight requests at equal memory. — [source](https://arxiv.org/html/2605.17170v1)
- The paper's stated limitations are per-workload calibration, reliance on chat-template markers, and INT2/INT4-only allocation. — [source](https://arxiv.org/html/2605.17170v1)
