<!-- llms-explorer concept facts · https://llms-explorer.com/tree/mtmd-mergeable-bitmaps-and-temporal-frame-mergin/ · pack 2026-10-05 · ~432 tokens -->

# mtmd mergeable bitmaps and temporal frame merging for Qwen-VL video

> Existing dossiers hold the mergeable flag, `n_merge_frames` from `clip_model_n_temporal_merge` capped at 2, and `mtmd_group_mergeable_bitmaps`; this file adds only header-level details.

Parent: [Mac local LLMs: llama.cpp internals](https://llms-explorer.com/tree/mac-local-llms-llama-cpp-internals/) · 1 facets · 5 facts · page: https://llms-explorer.com/tree/mtmd-mergeable-bitmaps-and-temporal-frame-mergin/

## Facts

- Existing dossiers hold the mergeable flag, `n_merge_frames` from `clip_model_n_temporal_merge` capped at 2, and `mtmd_group_mergeable_bitmaps`; this file adds only header-level details. — source: `asserted`
- The `mtmd_bitmap` header comment says `mtmd_tokenize()` performs the merging for Qwen-VL style models but the caller must call `mtmd_bitmap_set_mergeable(true)` on every frame. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/mtmd/mtmd.h)
- A bitmap created with null data is a placeholder for counting tokens: pass it through `mtmd_tokenize()` then call `mtmd_*_get_n_tokens()`; passing a placeholder to `mtmd_encode()` returns an error. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/mtmd/mtmd.h)
- The header comment on `mtmd_bitmap_set_mergeable` says a mergeable bitmap can be temporally merged with an adjacent mergeable bitmap only by certain video input models. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/mtmd/mtmd.h)
- The header comments that the temporal-position count equals max(t,h,w) for M-RoPE models and n_tokens otherwise. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/mtmd/mtmd.h)
