llama.cpp mtmd batching of consecutive image chunks into one non-causal decode
Parent: Mac local LLMs: llama.cpp internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
`mtmd_batch` is an encode-side batch: `mtmd_batch_encode` merges every entry's `batch_f32` into the first chunk's and runs one `mtmd_encode_chunk_impl`, and the server then decodes each chunk separately with `mtmd_helper_decode_image_chunk`.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- `mtmd_batch` is an encode-side batch: `mtmd_batch_encode` merges every entry's `batch_f32` into the first chunk's and runs one `mtmd_encode_chunk_impl`, and the server then decodes each chunk separately with `mtmd_helper_decode_image_chunk`. [source]
- `mtmd_batch_add_chunk` returns 1 for a text chunk or an unsupported model, 2 when the model cannot batch and the batch already has an entry or when adding the chunk would exceed `batch_max_tokens`, and 3 when the chunk cannot be batched with the first entry. [source]
- Two image chunks are batchable only when `nx`, `ny` and `pos` are equal; audio batching is marked TODO. [source]
- `batch_max_tokens` defaults to 1024; the header says it is not a hard limit because the first image is always added. [source]
- `llama-server` exposes the limit as `--mtmd-batch-max-tokens N` (default 1024, env `LLAMA_ARG_MTMD_BATCH_MAX_TOKENS`) and copies it into `mparams.batch_max_tokens`. [source]
- `process_mtmd_chunk` in `server-context.cpp` creates a new batch when none exists or the old one is used up, adds the current chunk, loops over later media chunks until one is refused, logs `encoding mtmd batch from idx = %zu, n_chunks = %d` at trace level, encodes once, then decodes the current chunk from its slice. [source]
- The server's decode path passes a speculative-decoding callback so the draft context also processes the image embeddings. [source]
- `mtmd_decode_use_non_causal` returns true for Gemma 4 UV, Gemma 3 and DeepSeek 4 V projectors, returns true for Gemma 4 V unless the mmproj embedding width is 1536 or 2560 (the E2B and E4B models, which stay causal), and false otherwise. [source]
- `mtmd_get_memory_usage` returns `use_non_causal` and `image_max_tokens`, and the header says that for non-causal models `max_tokens` must not exceed `n_ubatch` of the llama context. [source]
- The helper `mtmd_helper_decode_image_chunk` splits an image's embeddings into `n_batch` slices and decodes slice by slice, logging `decoding %s batch %d/%d`. [source]
- `mtmd_batch_get_output_embd` finds a chunk's embeddings by summing earlier entries' `n_tokens * n_embd_out` and returns null if the chunk is not in the batch or the batch is not yet encoded. [source]
Children
- No children recorded.