<!-- llms-explorer concept facts · https://llms-explorer.com/tree/sparse-v-attention-gated-dequant/ · pack 2026-10-05 · ~191 tokens -->

# Sparse V attention-gated dequant

> Already covered: Sparse V skips value dequantization for positions whose softmax weight is below 1e-6, gives +22.8% decode at 32K on MoE (0.76x to 0.93x of q8_0), is Metal-only in the fork, and is disabled with TURBO_SPARSE_V=0.

Parent: [Mac local LLMs: KV cache sizing and quantization](https://llms-explorer.com/tree/mac-local-llms-kv-cache-sizing-and-quantization/) · 1 facets · 1 facts · page: https://llms-explorer.com/tree/sparse-v-attention-gated-dequant/

## Facts

- Already covered: Sparse V skips value dequantization for positions whose softmax weight is below 1e-6, gives +22.8% decode at 32K on MoE (0.76x to 0.93x of q8_0), is Metal-only in the fork, and is disabled with TURBO_SPARSE_V=0. — source: `asserted`
