Gemma 4 sliding-window layer exclusion from KV quantization
Parent: Mac local LLMs: KV cache sizing and quantization · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
vLLM documents that some attention layer types, e.g. sliding-window, are more sensitive to KV-cache quantization and provides `--kv-cache-dtype-skip-layers` to leave chosen layers at native dtype while the rest use the quantized dtype.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- vLLM documents that some attention layer types, e.g. sliding-window, are more sensitive to KV-cache quantization and provides `--kv-cache-dtype-skip-layers` to leave chosen layers at native dtype while the rest use the quantized dtype. [source]
- The vLLM skip flag accepts layer-type names (`sliding_window`) or explicit layer indices (e.g. `0 1 23`), and a `kv_cache_dtype_skip_layers=["sliding_window"]` Python argument. [source]
- A downstream port commit message cites upstream per-cache rotation for iSWA (ggml-org#21513), giving gpt-oss-20b separate rotation inputs (attn_inp_k_rot, attn_inp_v_rot) for its base and SWA caches. [source]
- Inference: because sliding-window layers hold only a short ring (512 or 1024 tokens), keeping them at fp16 costs little memory, so the exclusion rule trades almost nothing for the layers it protects. [source]
Corrections and disagreements
- CONTRADICTS kv-cache-quantization-tradeoffs-on-apple-gpus.md (rotation cannot apply to Gemma 4 mixed-precision cache): per-cache iSWA rotation exists upstream, so the attn_rot_k = 0 log in #21394 is a point-in-time build observation, not a design limit. [source]
Children
- No children recorded.