Sparse V attention-gated dequant
Parent: Mac local LLMs: KV cache sizing and quantization · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Already covered: Sparse V skips value dequantization for positions whose softmax weight is below 1e-6, gives +22.8% decode at 32K on MoE (0.76x to 0.93x of q8_0), is Metal-only in the fork, and is disabled with TURBO_SPARSE_V=0.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Already covered: Sparse V skips value dequantization for positions whose softmax weight is below 1e-6, gives +22.8% decode at 32K on MoE (0.76x to 0.93x of q8_0), is Metal-only in the fork, and is disabled with TURBO_SPARSE_V=0. [source]
Children
- No children recorded.