<!-- llms-explorer concept facts · https://llms-explorer.com/tree/mmproj-versioning-and-mismatch-errors-for-multim/ · pack 2026-10-05 · ~946 tokens -->

# mmproj versioning and mismatch errors for multimodal GGUFs

> Three distinct load failures exist: embedding-size mismatch (wrong model family), unknown projector type (runtime older than the file), and a loaded but unusable projector (web UI still says no vision).

Parent: [Mac local LLMs: llama.cpp internals](https://llms-explorer.com/tree/mac-local-llms-llama-cpp-internals/) · 1 facets · 15 facts · page: https://llms-explorer.com/tree/mmproj-versioning-and-mismatch-errors-for-multim/

## Facts

- Three distinct load failures exist: embedding-size mismatch (wrong model family), unknown projector type (runtime older than the file), and a loaded but unusable projector (web UI still says no vision). — source: `asserted`
- Jun 2026: ggml-org's gemma-4-12B-it-GGUF Q8_0 mmproj failed to load on a current CUDA image and on macOS Apple Silicon. — [source](https://huggingface.co/ggml-org/gemma-4-12B-it-GGUF/discussions/3)
- Unknown projector type: `load_hparams: unknown projector type: gemma4uv`, then `mtmd_init_from_file: error: Failed to load CLIP model from .../mmproj-gemma-4-12B-it-Q8_0.gguf`, then `srv load_model: failed to load multimodal model`. — [source](https://huggingface.co/ggml-org/gemma-4-12B-it-GGUF/discussions/3)
- The embedding mismatch error is followed by `hint: you may be using wrong mmproj`. — [source](https://github.com/ggml-org/llama.cpp/discussions/22190)
- An mmproj from a different family can load without error yet leave the web UI image control disabled with "Image processing requires a vision model". — [source](https://github.com/ggml-org/llama.cpp/discussions/22190)
- Which llama.cpp build first recognised `gemma4uv`; the cached discussion shows no maintainer reply or fix. — source: `asserted`
- A fresh llama.cpp can fail on a fresh mmproj with `unknown projector type: gemma4uv`; the file is newer than the runtime, so updating llama.cpp, not re-downloading, is the lead fix to try (inferred). — [source](https://huggingface.co/ggml-org/gemma-4-12B-it-GGUF/discussions/3)
- The same `gemma4uv` failure was reported on macOS Apple Silicon with `llama-server -hf ggml-org/gemma-4-12B-it-GGUF:Q8_0`, so it is not CUDA-specific. — [source](https://huggingface.co/ggml-org/gemma-4-12B-it-GGUF/discussions/3)
- `-hf` downloads the mmproj together with the model, so a stale local cache plus a new runtime (or the reverse) mismatches without the user choosing a file. — source: `asserted`
- The `n_embd` mismatch error (2816 text versus 1536 mmproj) carries the hint `you may be using wrong mmproj`. — [source](https://github.com/ggml-org/llama.cpp/discussions/22190)
- An InternVL3-2B mmproj passed to `--mmproj` with gemma3, gemma4, gpt-oss and qwen3.6 text models did not enable images in the web UI. — [source](https://github.com/ggml-org/llama.cpp/discussions/22190)
- A finetune of gemma4-31b worked with the mmproj from unsloth's gemma4 31b repo: the mmproj follows the base architecture, not the repo. — [source](https://github.com/ggml-org/llama.cpp/discussions/22190)
- A user with a working gemma4-26b mmproj still hit an out-of-memory crash on a 24 GB GPU with the smallest image, so projector memory is separate from a text-only fit. — [source](https://github.com/ggml-org/llama.cpp/discussions/22190)
- In router `--models-dir`, a multimodal model must sit in its own subdirectory and the projector file name must start with `mmproj` (example `mmproj-F16.gguf`); otherwise the router cannot pair it. — [source](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)
- In router mode `/v1/models` reports `architecture.input_modalities` (for example `["text","image"]`), which shows whether the projector was picked up. — [source](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)
