<!-- llms-explorer concept facts · https://llms-explorer.com/tree/gemma-4-attention-k-eq-v-global-layer-key-as-val/ · pack 2026-10-05 · ~719 tokens -->

# Gemma 4 attention_k_eq_v global-layer key-as-value sharing

> A converter or quantizer that expects a `v_proj` weight on every global layer will not find one when `attention_k_eq_v` is true; mlx-lm creates `v_proj` only when `not use_k_eq_v`.

Parent: [Mac local LLMs: KV cache sizing and quantization](https://llms-explorer.com/tree/mac-local-llms-kv-cache-sizing-and-quantization/) · 1 facets · 10 facts · page: https://llms-explorer.com/tree/gemma-4-attention-k-eq-v-global-layer-key-as-val/

## Facts

- A converter or quantizer that expects a `v_proj` weight on every global layer will not find one when `attention_k_eq_v` is true; mlx-lm creates `v_proj` only when `not use_k_eq_v`. — source: `asserted`
- In mlx-lm's Gemma 4 text attention, `use_k_eq_v` is `attention_k_eq_v and not is_sliding`, and a source comment says it applies to the 26B and 31B models. — [source](https://raw.githubusercontent.com/ml-explore/mlx-lm/main/mlx_lm/models/gemma4_text.py)
- mlx-lm sets `values = keys` straight after `k_proj`, before `k_norm` and RoPE, and only creates a separate `v_proj` when `use_k_eq_v` is false. — [source](https://raw.githubusercontent.com/ml-explore/mlx-lm/main/mlx_lm/models/gemma4_text.py)
- mlx-lm applies `v_norm` (`RMSNormNoScale`) to values and `k_norm` plus RoPE to keys, so cached K and V differ under k_eq_v. — [source](https://raw.githubusercontent.com/ml-explore/mlx-lm/main/mlx_lm/models/gemma4_text.py)
- mlx-lm uses `num_global_key_value_heads` as the K/V head count on k_eq_v layers when that config field is not None. — [source](https://raw.githubusercontent.com/ml-explore/mlx-lm/main/mlx_lm/models/gemma4_text.py)
- mlx-lm sets the Gemma 4 attention scale to 1.0. — [source](https://raw.githubusercontent.com/ml-explore/mlx-lm/main/mlx_lm/models/gemma4_text.py)
- Hugging Face transformers names the flag `use_alternative_attention = config.attention_k_eq_v and not is_sliding`, sets `v_proj` to None in that case, and uses `value_states = key_states` before `k_norm`. — [source](https://raw.githubusercontent.com/huggingface/transformers/main/src/transformers/models/gemma4/modeling_gemma4.py)
- Hugging Face applies `v_norm = Gemma4RMSNorm(head_dim, with_scale=False)` to values, separate from the scaled `k_norm` on keys. — [source](https://raw.githubusercontent.com/huggingface/transformers/main/src/transformers/models/gemma4/modeling_gemma4.py)
- In transformers, KV-shared layers read `shared_kv_states[layer_type]` and have no `k_proj`, `v_proj`, `k_norm` or `v_norm`; the last non-shared layer of each attention type stores its full-length K/V into that dict. — [source](https://raw.githubusercontent.com/huggingface/transformers/main/src/transformers/models/gemma4/modeling_gemma4.py)
- Both implementations set the attention scaling to 1.0, so Gemma 4 relies on q_norm and k_norm instead of 1/sqrt(d) scaling. — source: `asserted`
