Sparse V attention-gated dequant

Parent: Mac local LLMs: KV cache sizing and quantization · Published reference · snapshot 2026-10-05

↓ Facts as markdownall context files

Already covered: Sparse V skips value dequantization for positions whose softmax weight is below 1e-6, gives +22.8% decode at 32K on MoE (0.76x to 0.93x of q8_0), is Metal-only in the fork, and is disabled with TURBO_SPARSE_V=0.

These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.

Facts

Children

← the whole tree · 3D view· how to read this page