<!-- llms-explorer concept facts · https://llms-explorer.com/tree/llama-cpp-mtmd-batching-of-consecutive-image-chu/ · pack 2026-10-05 · ~931 tokens -->

# llama.cpp mtmd batching of consecutive image chunks into one non-causal decode

> `mtmd_batch` is an encode-side batch: `mtmd_batch_encode` merges every entry's `batch_f32` into the first chunk's and runs one `mtmd_encode_chunk_impl`, and the server then decodes each chunk separately with `mtmd_helper_decode_image_chunk`.

Parent: [Mac local LLMs: llama.cpp internals](https://llms-explorer.com/tree/mac-local-llms-llama-cpp-internals/) · 1 facets · 11 facts · page: https://llms-explorer.com/tree/llama-cpp-mtmd-batching-of-consecutive-image-chu/

## Facts

- `mtmd_batch` is an encode-side batch: `mtmd_batch_encode` merges every entry's `batch_f32` into the first chunk's and runs one `mtmd_encode_chunk_impl`, and the server then decodes each chunk separately with `mtmd_helper_decode_image_chunk`. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/mtmd/mtmd.cpp)
- `mtmd_batch_add_chunk` returns 1 for a text chunk or an unsupported model, 2 when the model cannot batch and the batch already has an entry or when adding the chunk would exceed `batch_max_tokens`, and 3 when the chunk cannot be batched with the first entry. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/mtmd/mtmd.cpp)
- Two image chunks are batchable only when `nx`, `ny` and `pos` are equal; audio batching is marked TODO. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/mtmd/mtmd.cpp)
- `batch_max_tokens` defaults to 1024; the header says it is not a hard limit because the first image is always added. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/mtmd/mtmd.h)
- `llama-server` exposes the limit as `--mtmd-batch-max-tokens N` (default 1024, env `LLAMA_ARG_MTMD_BATCH_MAX_TOKENS`) and copies it into `mparams.batch_max_tokens`. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/README.md)
- `process_mtmd_chunk` in `server-context.cpp` creates a new batch when none exists or the old one is used up, adds the current chunk, loops over later media chunks until one is refused, logs `encoding mtmd batch from idx = %zu, n_chunks = %d` at trace level, encodes once, then decodes the current chunk from its slice. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-context.cpp)
- The server's decode path passes a speculative-decoding callback so the draft context also processes the image embeddings. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-context.cpp)
- `mtmd_decode_use_non_causal` returns true for Gemma 4 UV, Gemma 3 and DeepSeek 4 V projectors, returns true for Gemma 4 V unless the mmproj embedding width is 1536 or 2560 (the E2B and E4B models, which stay causal), and false otherwise. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/mtmd/mtmd.cpp)
- `mtmd_get_memory_usage` returns `use_non_causal` and `image_max_tokens`, and the header says that for non-causal models `max_tokens` must not exceed `n_ubatch` of the llama context. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/mtmd/mtmd.h)
- The helper `mtmd_helper_decode_image_chunk` splits an image's embeddings into `n_batch` slices and decodes slice by slice, logging `decoding %s batch %d/%d`. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/mtmd/mtmd-helper.cpp)
- `mtmd_batch_get_output_embd` finds a chunk's embeddings by summing earlier entries' `n_tokens * n_embd_out` and returns null if the chunk is not in the batch or the batch is not yet encoded. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/mtmd/mtmd.cpp)
