mtmd mergeable bitmaps and temporal frame merging for Qwen-VL video
Parent: Mac local LLMs: llama.cpp internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Existing dossiers hold the mergeable flag, `n_merge_frames` from `clip_model_n_temporal_merge` capped at 2, and `mtmd_group_mergeable_bitmaps`; this file adds only header-level details.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Existing dossiers hold the mergeable flag, `n_merge_frames` from `clip_model_n_temporal_merge` capped at 2, and `mtmd_group_mergeable_bitmaps`; this file adds only header-level details. [source]
- The `mtmd_bitmap` header comment says `mtmd_tokenize()` performs the merging for Qwen-VL style models but the caller must call `mtmd_bitmap_set_mergeable(true)` on every frame. [source]
- A bitmap created with null data is a placeholder for counting tokens: pass it through `mtmd_tokenize()` then call `mtmd_*_get_n_tokens()`; passing a placeholder to `mtmd_encode()` returns an error. [source]
- The header comment on `mtmd_bitmap_set_mergeable` says a mergeable bitmap can be temporally merged with an adjacent mergeable bitmap only by certain video input models. [source]
- The header comments that the temporal-position count equals max(t,h,w) for M-RoPE models and n_tokens otherwise. [source]
Children
- No children recorded.