<!-- llms-explorer concept facts · https://llms-explorer.com/tree/gemma-4-sliding-window-layer-exclusion-from-kv-q/ · pack 2026-10-05 · ~486 tokens -->

# Gemma 4 sliding-window layer exclusion from KV quantization

> vLLM documents that some attention layer types, e.g. sliding-window, are more sensitive to KV-cache quantization and provides `--kv-cache-dtype-skip-layers` to leave chosen layers at native dtype while the rest use the quantized dtype.

Parent: [Mac local LLMs: KV cache sizing and quantization](https://llms-explorer.com/tree/mac-local-llms-kv-cache-sizing-and-quantization/) · 2 facets · 5 facts · page: https://llms-explorer.com/tree/gemma-4-sliding-window-layer-exclusion-from-kv-q/

## Facts

- vLLM documents that some attention layer types, e.g. sliding-window, are more sensitive to KV-cache quantization and provides `--kv-cache-dtype-skip-layers` to leave chosen layers at native dtype while the rest use the quantized dtype. — [source](https://docs.vllm.ai/en/latest/features/quantization/quantized_kvcache/)
- The vLLM skip flag accepts layer-type names (`sliding_window`) or explicit layer indices (e.g. `0 1 23`), and a `kv_cache_dtype_skip_layers=["sliding_window"]` Python argument. — [source](https://docs.vllm.ai/en/latest/features/quantization/quantized_kvcache/)
- A downstream port commit message cites upstream per-cache rotation for iSWA (ggml-org#21513), giving gpt-oss-20b separate rotation inputs (attn_inp_k_rot, attn_inp_v_rot) for its base and SWA caches. — [source](https://github.com/ggml-org/llama.cpp/pull/21038)
- Inference: because sliding-window layers hold only a short ring (512 or 1024 tokens), keeping them at fp16 costs little memory, so the exclusion rule trades almost nothing for the layers it protects. — source: `asserted`

## Corrections and disagreements

- CONTRADICTS kv-cache-quantization-tradeoffs-on-apple-gpus.md (rotation cannot apply to Gemma 4 mixed-precision cache): per-cache iSWA rotation exists upstream, so the attn_rot_k = 0 log in #21394 is a point-in-time build observation, not a design limit. — source: `asserted`
