TriAxialKV per-role mixed-precision KV quantization for agentic prefills
Parent: Mac local LLMs: KV cache sizing and quantization · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
TriAxialKV tags tokens on three axes with 4 temporal values, 2 modal values and 7 semantic values, and combinations that never occur drop out of the optimization.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- TriAxialKV tags tokens on three axes with 4 temporal values, 2 modal values and 7 semantic values, and combinations that never occur drop out of the optimization. [source]
- The tagger runs on the CPU at scheduling time from chat-template special tokens, needing no model inference. [source]
- Calibration uses 5% of a dataset and hooks on four to six layers spread across depth, capturing prefill Q, K and V only. [source]
- Per-tag distortion D_k(b) is attention-output MSE, aggregated as max over heads, mean over requests, sum over layers; the paper argues output MSE is not monotonic in raw KV error. [source]
- The allocator solves INT2 or INT4 per tag under an average-bit budget, by exhaustive search for 22 or fewer tags and a greedy per-bit-gain rule otherwise. [source]
- The quantizer is asymmetric groupwise with group size 32, INT4 per-token for K and V, INT2 per-channel for keys and per-token for values. [source]
- Leftover INT2-key tokens that do not fill a group of 32 are routed through the INT4 per-token path rather than a full-precision residual buffer of 128 tokens per layer. [source]
- INT2 and INT4 page pools share one virtual address space split at an offset derived from the calibrated average bitwidth. [source]
- The decode path partitions each request's page table so INT2 pointers precede INT4, giving bitwidth-homogeneous flash-decoding splits merged by online log-sum-exp. [source]
- Reported throughput is 1.26x BF16 on Qwen3-VL-8B (B200), 1.32x on Qwen3-VL-32B (B200) and 1.52x on Qwen3-VL-32B (H100), with 3.4 to 4.0x more concurrent in-flight requests at equal memory. [source]
- The paper's stated limitations are per-workload calibration, reliance on chat-template markers, and INT2/INT4-only allocation. [source]
Children
- No children recorded.