<!-- llms-explorer concept facts · https://llms-explorer.com/tree/gemma4clippablelinear-input-output-clamp-scalars/ · pack 2026-10-05 · ~550 tokens -->

# Gemma4ClippableLinear input/output clamp scalars in the mmproj

> Clamping bounds each wrapped linear's input and output; it does not bound the residual stream between layers.

Parent: [Mac local LLMs: MLX kernels, numerics and internals](https://llms-explorer.com/tree/mac-local-llms-mlx-kernels-numerics-and-internals/) · 1 facets · 7 facts · page: https://llms-explorer.com/tree/gemma4clippablelinear-input-output-clamp-scalars/

## Facts

- Clamping bounds each wrapped linear's input and output; it does not bound the residual stream between layers. — source: `asserted`
- `Gemma4ClippableLinear` holds a bias-free `nn.Linear` and, when `use_clipped_linears` is set, buffers `input_min`, `input_max`, `output_min` and `output_max` initialised to -inf and +inf. — [source](https://raw.githubusercontent.com/huggingface/transformers/main/src/transformers/models/gemma4/modeling_gemma4.py)
- Its forward pass applies `torch.clamp` to the input, then the linear, then `torch.clamp` to the output, each only when `use_clipped_linears` is true. — [source](https://raw.githubusercontent.com/huggingface/transformers/main/src/transformers/models/gemma4/modeling_gemma4.py)
- The audio tower uses it for attention q, k, v and post projections, the feed-forward `ffw_layer_1` and `ffw_layer_2`, and the light-conv `linear_start` and `linear_end`. — [source](https://raw.githubusercontent.com/huggingface/transformers/main/src/transformers/models/gemma4/modeling_gemma4.py)
- The vision tower uses it for the MLP `gate_proj`, `up_proj` and `down_proj` and for attention `q_proj`, `k_proj`, `v_proj` and `o_proj`. — [source](https://raw.githubusercontent.com/huggingface/transformers/main/src/transformers/models/gemma4/modeling_gemma4.py)
- The Gemma 4 text attention builds `q_proj`, `k_proj`, `v_proj` and `o_proj` as plain `nn.Linear`, not the clippable class. — [source](https://raw.githubusercontent.com/huggingface/transformers/main/src/transformers/models/gemma4/modeling_gemma4.py)
- A GGUF converter must keep the 4 scalars per clippable linear in F32 or special-case them, because they are per-layer scalars that exist for each of the vision and audio linears listed above. — source: `asserted`
