Gemma 4 attention_k_eq_v global-layer key-as-value sharing
Parent: Mac local LLMs: KV cache sizing and quantization · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
A converter or quantizer that expects a `v_proj` weight on every global layer will not find one when `attention_k_eq_v` is true; mlx-lm creates `v_proj` only when `not use_k_eq_v`.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- A converter or quantizer that expects a `v_proj` weight on every global layer will not find one when `attention_k_eq_v` is true; mlx-lm creates `v_proj` only when `not use_k_eq_v`. [source]
- In mlx-lm's Gemma 4 text attention, `use_k_eq_v` is `attention_k_eq_v and not is_sliding`, and a source comment says it applies to the 26B and 31B models. [source]
- mlx-lm sets `values = keys` straight after `k_proj`, before `k_norm` and RoPE, and only creates a separate `v_proj` when `use_k_eq_v` is false. [source]
- mlx-lm applies `v_norm` (`RMSNormNoScale`) to values and `k_norm` plus RoPE to keys, so cached K and V differ under k_eq_v. [source]
- mlx-lm uses `num_global_key_value_heads` as the K/V head count on k_eq_v layers when that config field is not None. [source]
- mlx-lm sets the Gemma 4 attention scale to 1.0. [source]
- Hugging Face transformers names the flag `use_alternative_attention = config.attention_k_eq_v and not is_sliding`, sets `v_proj` to None in that case, and uses `value_states = key_states` before `k_norm`. [source]
- Hugging Face applies `v_norm = Gemma4RMSNorm(head_dim, with_scale=False)` to values, separate from the scaled `k_norm` on keys. [source]
- In transformers, KV-shared layers read `shared_kv_states[layer_type]` and have no `k_proj`, `v_proj`, `k_norm` or `v_norm`; the last non-shared layer of each attention type stores its full-length K/V into that dict. [source]
- Both implementations set the attention scaling to 1.0, so Gemma 4 relies on q_norm and k_norm instead of 1/sqrt(d) scaling. [source]
Children
- No children recorded.