<!-- llms-explorer concept facts · https://llms-explorer.com/tree/llama-cpp-mtmd-batch-max-tokens-tuning-against-n/ · pack 2026-10-05 · ~279 tokens -->

# llama.cpp --mtmd-batch-max-tokens tuning against n_ubatch for non-causal project

> Already covered: llama-cpp-mtmd-batching-of-consecutive-image-chu.md holds the default 1024, the not-a-hard-limit first-image rule, the `--mtmd-batch-max-tokens` flag, and the rule that for non-causal models `max_tokens` must not exceed `n_ubatch`. The header comment on `mtmd_memory_usage` and th...

Parent: [Mac local LLMs: llama.cpp internals](https://llms-explorer.com/tree/mac-local-llms-llama-cpp-internals/) · 1 facets · 1 facts · page: https://llms-explorer.com/tree/llama-cpp-mtmd-batch-max-tokens-tuning-against-n/

## Facts

- Already covered: llama-cpp-mtmd-batching-of-consecutive-image-chu.md holds the default 1024, the not-a-hard-limit first-image rule, the `--mtmd-batch-max-tokens` flag, and the rule that for non-causal models `max_tokens` must not exceed `n_ubatch`. The header comment on `mtmd_memory_usage` and the `batch_max_tokens` comment add nothing beyond that; no source gives a tuning guideline relating the two. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/mtmd/mtmd.h)
