Gemma4ClippableLinear input/output clamp scalars in the mmproj
Parent: Mac local LLMs: MLX kernels, numerics and internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Clamping bounds each wrapped linear's input and output; it does not bound the residual stream between layers.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Clamping bounds each wrapped linear's input and output; it does not bound the residual stream between layers. [source]
- `Gemma4ClippableLinear` holds a bias-free `nn.Linear` and, when `use_clipped_linears` is set, buffers `input_min`, `input_max`, `output_min` and `output_max` initialised to -inf and +inf. [source]
- Its forward pass applies `torch.clamp` to the input, then the linear, then `torch.clamp` to the output, each only when `use_clipped_linears` is true. [source]
- The audio tower uses it for attention q, k, v and post projections, the feed-forward `ffw_layer_1` and `ffw_layer_2`, and the light-conv `linear_start` and `linear_end`. [source]
- The vision tower uses it for the MLP `gate_proj`, `up_proj` and `down_proj` and for attention `q_proj`, `k_proj`, `v_proj` and `o_proj`. [source]
- The Gemma 4 text attention builds `q_proj`, `k_proj`, `v_proj` and `o_proj` as plain `nn.Linear`, not the clippable class. [source]
- A GGUF converter must keep the 4 scalars per clippable linear in F32 or special-case them, because they are per-layer scalars that exist for each of the vision and audio linears listed above. [source]
Children
- No children recorded.