<!-- llms-explorer concept facts · https://llms-explorer.com/tree/mtmd-tokenize-from-parts-per-segment-tokenizatio/ · pack 2026-10-05 · ~796 tokens -->

# mtmd_tokenize_from_parts per-segment tokenization for marked input

> `mtmd_input_part` holds a `const mtmd_input_text *` and a `const mtmd_bitmap *`, with the comment that only one of them can be set.

Parent: [Mac local LLMs: llama.cpp internals](https://llms-explorer.com/tree/mac-local-llms-llama-cpp-internals/) · 1 facets · 10 facts · page: https://llms-explorer.com/tree/mtmd-tokenize-from-parts-per-segment-tokenizatio/

## Facts

- `mtmd_input_part` holds a `const mtmd_input_text *` and a `const mtmd_bitmap *`, with the comment that only one of them can be set. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/mtmd/mtmd.h)
- The header documents two uses for `mtmd_tokenize_from_parts`: when you do not want media markers (markers are tokenized as normal text) and when you want to control `parse_special` for each text part. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/mtmd/mtmd.h)
- The header says the per-part `add_special` is ignored; the single `add_special` argument of the function applies. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/mtmd/mtmd.h)
- `mtmd_tokenize_from_parts` returns 1 if a part has both text and bitmap set or neither, or if a text part's text pointer is null, and returns 2 when the tokenizer throws. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/mtmd/mtmd.cpp)
- The parts constructor copies each text part with its own `parse_special` and sets the tokenizer-level `parse_special` to true, used only for text returned by lazy bitmaps. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/mtmd/mtmd.cpp)
- `mtmd_tokenize` returns 0 on success, 1 when the number of bitmaps does not match the number of markers, and 2 on media preprocessing error; the marker-count mismatch is a thrown `number of media markers in text (%zu) does not match number of bitmaps (%zu)`. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/mtmd/mtmd.cpp)
- A lazy bitmap's callback is called with an increasing chunk index until it returns -1 (EOF) or -2 (error); it may return a bitmap or a heap-allocated text string, not both, and not another lazy bitmap. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/mtmd/mtmd.cpp)
- `tokenize` computes `n_merge_frames` from `clip_model_n_temporal_merge`, asserts it is at most 2, and groups bitmaps flagged mergeable with `mtmd_group_mergeable_bitmaps` before adding media. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/mtmd/mtmd.cpp)
- When `add_special` is true and the vocab adds BOS, BOS goes at the start of the first text chunk, or a new BOS-only text chunk is inserted at the front when the first chunk is media; the same block checks the vocab's add-EOS flag. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/mtmd/mtmd.cpp)
- An `mtmd_bitmap` can carry an optional ID for KV cache tracking, and a lazy bitmap is meant to be tracked by one ID such as the file hash. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/mtmd/mtmd.h)
