<!-- llms-explorer concept facts · https://llms-explorer.com/tree/llama-cpp-mtmd-video-and-webp-input-via-ffmpeg-s/ · pack 2026-10-05 · ~907 tokens -->

# llama.cpp mtmd video and webp input via ffmpeg subprocess (MTMD_VIDEO, PR 27520)

> PR 27520 "mtmd: support webp via ffmpeg" merged on 2026-08-21 at 23:38 UTC as commit d775b89, fixing issues 27443 and 12410 and superseding PR 24217.

Parent: [Mac local LLMs: llama.cpp internals](https://llms-explorer.com/tree/mac-local-llms-llama-cpp-internals/) · 1 facets · 11 facts · page: https://llms-explorer.com/tree/llama-cpp-mtmd-video-and-webp-input-via-ffmpeg-s/

## Facts

- PR 27520 "mtmd: support webp via ffmpeg" merged on 2026-08-21 at 23:38 UTC as commit d775b89, fixing issues 27443 and 12410 and superseding PR 24217. — [source](https://github.com/ggml-org/llama.cpp/pull/27520)
- The PR body argues the ffmpeg route needs no vendored cpp or h files, no extra compile time and no binary bloat for a feature under 1% of users would use. — [source](https://github.com/ggml-org/llama.cpp/pull/27520)
- PR 27520 changes `tools/mtmd/mtmd-helper.cpp` and `tools/mtmd/mtmd-helper.h`, adds `is_webp_file` and `decode_webp_with_ffmpeg`, and adds no new public API. — [source](https://github.com/ggml-org/llama.cpp/pull/27520)
- `is_webp_file` bounds-checks `len >= 12` and compares `RIFF` at offset 0 and `WEBP` at offset 8; the buffer is then forwarded to the ffmpeg subprocess as the video path does. — [source](https://github.com/ggml-org/llama.cpp/pull/27520)
- An automated review on the PR notes that animated webp share the `RIFF....WEBP` header, so `decode_webp_with_ffmpeg` reads one frame and the file never reaches the later "load as video" path; the review text says it may contain mistakes. — [source](https://github.com/ggml-org/llama.cpp/pull/27520)
- The same review notes `decode_webp_with_ffmpeg` calls `probe(0.0f)` and `probe` returns false when `orig_fps <= 0.0f`, so a still webp that ffprobe reports as `r_frame_rate=0/0` could fail with `failed to decode webp buffer`; this is stated as a risk, not a reproduced failure. — [source](https://github.com/ggml-org/llama.cpp/pull/27520)
- The same review notes `is_webp_file` is defined outside the `#ifdef MTMD_VIDEO` block but used only inside it, which gives an unused-function warning when `MTMD_VIDEO` is off; `tools/mtmd/CMakeLists.txt` turns `MTMD_VIDEO` off whenever `LLAMA_SUBPROCESS` is off. — [source](https://github.com/ggml-org/llama.cpp/pull/27520)
- A maintainer-side comment on the PR says using ffmpeg is a better route than the other options proposed for webp. — [source](https://github.com/ggml-org/llama.cpp/pull/27520)
- `llama-server` has `--video-fps N` (default 4.0), `--video-timestamp-interval N` (milliseconds between text timestamps, default 5000) and `--video-ffmpeg-dir DIR` (directory holding ffmpeg and ffprobe, default search in `PATH`), each with a `LLAMA_ARG_` environment variable. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/README.md)
- The chat API accepts `type: "input_video"` with `input_video.data` or `input_video.url` (remote URL, raw base64 or local path), accepts formats ffmpeg supports, and a local file needs `--media-path` and a `file://` prefix; `image_url` accepts formats stb_image reads, such as jpeg, png, tga, bmp and gif. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/README.md)
- `mtmd_bitmap_init_lazy` lets a video be read frame by frame without holding it in memory and be tracked by a single ID such as the file hash, and Qwen-VL style models merge consecutive bitmaps flagged with `mtmd_bitmap_set_mergeable(true)` into one chunk. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/mtmd/mtmd.h)
