llama.cpp gemma4uv projector type support and first supporting build
Parent: Mac local LLMs: llama.cpp internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
`PROJECTOR_TYPE_GEMMA4V`, `GEMMA4A`, `GEMMA4UV` and `GEMMA4UA` are mapped to the strings `gemma4v`, `gemma4a`, `gemma4uv` and `gemma4ua`.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- `PROJECTOR_TYPE_GEMMA4V`, `GEMMA4A`, `GEMMA4UV` and `GEMMA4UA` are mapped to the strings `gemma4v`, `gemma4a`, `gemma4uv` and `gemma4ua`. [source]
- Issue 24251 (opened 2026-06-06) reports a `gemma4uv` mmproj with 11 tensors, vision plus audio encoder, used with master commit 6b80c74f285390368b3c99c5e750f19e9b096e98 (2026-06-07). [source]
- Issue 24251 reports NaN logits and a sampler assertion failure with the GPU backend enabled on an RTX 5070 Ti with CUDA 13.0, and was closed as not planned. [source]
- Issue 24146 (opened 2026-06-04, build b9518) reports Gemma 4 12B image input returning repeated `<unused49>` tokens or empty output with the ggml-org mmproj on Windows and CUDA. [source]
- Issue 24146 diagnoses two causes: the resize never fills the soft-token budget (a 512x512 image became 528x528, 121 tokens, unlike the reference `Gemma4UnifiedImageProcessor`) and an F16 overflow. [source]
- A fork commit titled "Fix Gemma 4 12B vision: wrong patch count + F16 overflow (issue #24146)" is linked from issue 24146, and a commenter reported clean results after applying it. [source]
- PR 26488 ("mtmd: rebase Gemma 4 unified-vision fixes onto master") was a draft closed on 2026-08-03. [source]
- PR 27594 (ngxson, merged 2026-08-23) removes all non-pillow resize variants and corrects `resize_algo` for every model's config. [source]
- PR 28335 makes E2B and E4B always use causal attention, keeps global layers causal to match transformers, and fixes the vision token budget, superseding PR 24014. [source]
- PR 28335's model table lists gemma-4-12B-it as `use_bidirectional_attention: "vision"`, hidden size 3840, no PLE, "unified (encoder-free)"; 26B-A4B (hidden 2816) and 31B (hidden 5376) are also bidirectional; E2B (1536) and E4B (2560) are causal. [source]
- PR 28335 lists a 4-layer MTP drafter "gemma-4-31B-it-assistant" with width 1024 projecting from the 31B backbone (5376). [source]
- PR 28751 removes `sched_need_reserve = true` from `llama_context::set_causal_attn()` because the flag only changes KQ mask values, and fixes the `qwen4exp` graph that depended on it. [source]
- PR 28751 says the causal flag is flipped twice around each non-causal image chunk for Gemma 3, Gemma 4 and DeepSeek-V4, costing two scheduler reserves per image. [source]
- PR 28751's measured speedups on gemma-4-26B-A4B Q4_0 with a BF16 mmproj range from 1.27x (one image, H200) to 7.61x (24 images, `-c 32768 -ub 2048`, RTX 4090); no Apple silicon numbers are given. [source]
- The pull-request search for `gemma4uv` lists PR 28919 open (2026-09-15), PR 28751 merged, PR 28335 merged, PR 27594 merged, PR 26488 closed, and PR 23398 (Gemma4 MTP) merged 2026-06-07. [source]
- The issue search for `gemma4uv` lists issue 26981 (CUDA SIGABRT in `mtmd_helper_decode_image_chunk`, closed 2026-09-26) and issue 29686 (closed 2026-09-30). [source]
- Ollama's 2026-06-04 bump to b9509 states it carries the upstream Gemma 4 12B projector fixes for the `n_head=0` crash (issues 16479 and 16489). [source]
- Ollama's compat clip allowlist names `gemma4` but not a separate unified entry. [source]
- Upper bound: the first llama.cpp build to recognise `gemma4uv` is at or before b9509 (2026-06-04); the exact first PR was not found. [source]
- Every Ollama release pinning b10760 or later recognises `gemma4uv` because the pins are far above b9509. [source]
Children
- No children recorded.