mmproj versioning and mismatch errors for multimodal GGUFs
Parent: Mac local LLMs: llama.cpp internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Three distinct load failures exist: embedding-size mismatch (wrong model family), unknown projector type (runtime older than the file), and a loaded but unusable projector (web UI still says no vision).
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Three distinct load failures exist: embedding-size mismatch (wrong model family), unknown projector type (runtime older than the file), and a loaded but unusable projector (web UI still says no vision). [source]
- Jun 2026: ggml-org's gemma-4-12B-it-GGUF Q8_0 mmproj failed to load on a current CUDA image and on macOS Apple Silicon. [source]
- Unknown projector type: `load_hparams: unknown projector type: gemma4uv`, then `mtmd_init_from_file: error: Failed to load CLIP model from .../mmproj-gemma-4-12B-it-Q8_0.gguf`, then `srv load_model: failed to load multimodal model`. [source]
- The embedding mismatch error is followed by `hint: you may be using wrong mmproj`. [source]
- An mmproj from a different family can load without error yet leave the web UI image control disabled with "Image processing requires a vision model". [source]
- Which llama.cpp build first recognised `gemma4uv`; the cached discussion shows no maintainer reply or fix. [source]
- A fresh llama.cpp can fail on a fresh mmproj with `unknown projector type: gemma4uv`; the file is newer than the runtime, so updating llama.cpp, not re-downloading, is the lead fix to try (inferred). [source]
- The same `gemma4uv` failure was reported on macOS Apple Silicon with `llama-server -hf ggml-org/gemma-4-12B-it-GGUF:Q8_0`, so it is not CUDA-specific. [source]
- `-hf` downloads the mmproj together with the model, so a stale local cache plus a new runtime (or the reverse) mismatches without the user choosing a file. [source]
- The `n_embd` mismatch error (2816 text versus 1536 mmproj) carries the hint `you may be using wrong mmproj`. [source]
- An InternVL3-2B mmproj passed to `--mmproj` with gemma3, gemma4, gpt-oss and qwen3.6 text models did not enable images in the web UI. [source]
- A finetune of gemma4-31b worked with the mmproj from unsloth's gemma4 31b repo: the mmproj follows the base architecture, not the repo. [source]
- A user with a working gemma4-26b mmproj still hit an out-of-memory crash on a 24 GB GPU with the smallest image, so projector memory is separate from a text-only fit. [source]
- In router `--models-dir`, a multimodal model must sit in its own subdirectory and the projector file name must start with `mmproj` (example `mmproj-F16.gguf`); otherwise the router cannot pair it. [source]
- In router mode `/v1/models` reports `architecture.input_modalities` (for example `["text","image"]`), which shows whether the projector was picked up. [source]
Children
- No children recorded.