<!-- llms-explorer concept facts · https://llms-explorer.com/tree/llama-cpp-causal-attn-flag-scheduler-re-reserve/ · pack 2026-10-05 · ~1247 tokens -->

# llama.cpp causal_attn flag scheduler re-reserve cost per image (PR 28751)

> PR 28751 was opened on 2026-09-11 and merged on 2026-09-28 as commit ed7ac35e1ee49cb70e4dfa9f0a2ce39b0a5ec4ea after 14 checks passed.

Parent: [Mac local LLMs: llama.cpp internals](https://llms-explorer.com/tree/mac-local-llms-llama-cpp-internals/) · 1 facets · 19 facts · page: https://llms-explorer.com/tree/llama-cpp-causal-attn-flag-scheduler-re-reserve/

## Facts

- PR 28751 was opened on 2026-09-11 and merged on 2026-09-28 as commit ed7ac35e1ee49cb70e4dfa9f0a2ce39b0a5ec4ea after 14 checks passed. — [source](https://github.com/ggml-org/llama.cpp/pull/28751)
- Commit ed7ac35 changes `src/llama-context.cpp`, `src/llama-memory-hybrid-idx.cpp`, `src/llama-memory-hybrid-idx.h` and `src/models/qwen4exp.cpp`. — [source](https://github.com/ggml-org/llama.cpp/commit/ed7ac35e1ee49cb70e4dfa9f0a2ce39b0a5ec4ea)
- In `set_causal_attn`, ed7ac35 replaces `sched_need_reserve = true;` with a comment and a commented-out copy of the line. — [source](https://github.com/ggml-org/llama.cpp/commit/ed7ac35e1ee49cb70e4dfa9f0a2ce39b0a5ec4ea)
- The PR author found by grepping `cparams.causal_attn` in `src/` that the only read that changes tensor sizes is in `src/models/qwen4exp.cpp` lines 552-569, where it affects `blk_bias` and the size of `qsa->bias`. — [source](https://github.com/ggml-org/llama.cpp/pull/28751)
- am17an noted that CI runs with `GGML_SCHED_NO_REALLOC`, so a model whose graph shape depends on the flag would fail those tests. — [source](https://github.com/ggml-org/llama.cpp/pull/28751)
- The author's `test-llama-archs` toggle test showed that master plus only the first commit aborts qwen4exp with `ggml_backend_sched_alloc_splits: unexpected graph reallocation`, and the full PR passes all 366 tests with `-DGGML_SCHED_NO_REALLOC=ON`. — [source](https://github.com/ggml-org/llama.cpp/pull/28751)
- ggerganov wrote that disabling the `causal_attn` reserve (as in PR 28927) is fine, preferred a rewrite of the Qwen4 inference graph and memory over more changes, and then asked to keep the qwen4exp changes so CI does not break. — [source](https://github.com/ggml-org/llama.cpp/pull/28751)
- ddh0 reported on DeepSeek-V4-Flash-Vision-Exp that the graph was re-reserved before the first image, between every image and once to return to text generation, each taking about 5 seconds on their machine. — [source](https://github.com/ggml-org/llama.cpp/pull/28751)
- PR 28927 (dkrisman, opened 2026-09-15) made the identical change and was closed by ggerganov on 2026-09-28, the day PR 28751 merged. — [source](https://github.com/ggml-org/llama.cpp/pull/28927)
- PR 28927 measured each reserve at about 700 ms with `n_ctx = 262144` and two slots, and about 330 ms with `n_ctx = 131072` (allocating a 612 MiB CUDA-host compute buffer), with GPU utilization near 3% and one CPU core saturated. — [source](https://github.com/ggml-org/llama.cpp/pull/28927)
- PR 28927 reports on an RTX PRO 6000 Blackwell with Gemma 4 31B Q4_K_M: a 10-frame clip at `n_ctx = 131072` went from 10.46 s to 3.18 s, a single-image embedding decode from 405 ms to 66 ms, and a 280-frame clip at `n_ctx = 262144` from 489 s to 118 s. — [source](https://github.com/ggml-org/llama.cpp/pull/28927)
- PR 28927 states that `tools/mtmd/mtmd-helper.cpp::scope_non_causal` disables causal attention before every image chunk and restores it afterward, and that audio chunks use the same path. — [source](https://github.com/ggml-org/llama.cpp/pull/28927)
- PR 28927 states that `llama_context::encode()` changes `cparams.causal_attn` around `process_ubatch()` without requesting a scheduler reserve. — [source](https://github.com/ggml-org/llama.cpp/pull/28927)
- PR 28927 lists as untested: audio chunks, M-RoPE vision models such as Qwen VL, CPU-only backends and `--no-mmproj-offload`. — [source](https://github.com/ggml-org/llama.cpp/pull/28927)
- Tag b11232 was released on 2026-09-28 and is built from commit 6f767fe ("ggml-cpu: enable tiled flash attention for non-vector-multiple head dims on x86"). — [source](https://github.com/ggml-org/llama.cpp/releases/tag/b11232)
- A GitHub compare from ed7ac35 to b11232 lists 5 commits and 15 changed files, which means b11232 is ahead of the PR 28751 merge commit. — [source](https://github.com/ggml-org/llama.cpp/compare/ed7ac35e1ee49cb70e4dfa9f0a2ce39b0a5ec4ea...b11232)
- A GitHub compare from ed7ac35 to b11081 reports nothing to compare because ed7ac35 already contains all commits from b11081, so b11081 predates the merge. — [source](https://github.com/ggml-org/llama.cpp/compare/ed7ac35e1ee49cb70e4dfa9f0a2ce39b0a5ec4ea...b11081)
- Ollama tags v0.35.0 (pin b11081) lack the fix; v0.35.1 (pin b11232, see llama-cpp-bundled-revision-inside-ollama-release.md) includes it. — source: `asserted`
- A Mac user sending several images per request to an Ollama build older than v0.35.1 pays two scheduler reserves per image; the cost is larger with a large `num_ctx`. — source: `asserted`
