llama.cpp causal_attn flag scheduler re-reserve cost per image (PR 28751)
Parent: Mac local LLMs: llama.cpp internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
PR 28751 was opened on 2026-09-11 and merged on 2026-09-28 as commit ed7ac35e1ee49cb70e4dfa9f0a2ce39b0a5ec4ea after 14 checks passed.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- PR 28751 was opened on 2026-09-11 and merged on 2026-09-28 as commit ed7ac35e1ee49cb70e4dfa9f0a2ce39b0a5ec4ea after 14 checks passed. [source]
- Commit ed7ac35 changes `src/llama-context.cpp`, `src/llama-memory-hybrid-idx.cpp`, `src/llama-memory-hybrid-idx.h` and `src/models/qwen4exp.cpp`. [source]
- In `set_causal_attn`, ed7ac35 replaces `sched_need_reserve = true;` with a comment and a commented-out copy of the line. [source]
- The PR author found by grepping `cparams.causal_attn` in `src/` that the only read that changes tensor sizes is in `src/models/qwen4exp.cpp` lines 552-569, where it affects `blk_bias` and the size of `qsa->bias`. [source]
- am17an noted that CI runs with `GGML_SCHED_NO_REALLOC`, so a model whose graph shape depends on the flag would fail those tests. [source]
- The author's `test-llama-archs` toggle test showed that master plus only the first commit aborts qwen4exp with `ggml_backend_sched_alloc_splits: unexpected graph reallocation`, and the full PR passes all 366 tests with `-DGGML_SCHED_NO_REALLOC=ON`. [source]
- ggerganov wrote that disabling the `causal_attn` reserve (as in PR 28927) is fine, preferred a rewrite of the Qwen4 inference graph and memory over more changes, and then asked to keep the qwen4exp changes so CI does not break. [source]
- ddh0 reported on DeepSeek-V4-Flash-Vision-Exp that the graph was re-reserved before the first image, between every image and once to return to text generation, each taking about 5 seconds on their machine. [source]
- PR 28927 (dkrisman, opened 2026-09-15) made the identical change and was closed by ggerganov on 2026-09-28, the day PR 28751 merged. [source]
- PR 28927 measured each reserve at about 700 ms with `n_ctx = 262144` and two slots, and about 330 ms with `n_ctx = 131072` (allocating a 612 MiB CUDA-host compute buffer), with GPU utilization near 3% and one CPU core saturated. [source]
- PR 28927 reports on an RTX PRO 6000 Blackwell with Gemma 4 31B Q4_K_M: a 10-frame clip at `n_ctx = 131072` went from 10.46 s to 3.18 s, a single-image embedding decode from 405 ms to 66 ms, and a 280-frame clip at `n_ctx = 262144` from 489 s to 118 s. [source]
- PR 28927 states that `tools/mtmd/mtmd-helper.cpp::scope_non_causal` disables causal attention before every image chunk and restores it afterward, and that audio chunks use the same path. [source]
- PR 28927 states that `llama_context::encode()` changes `cparams.causal_attn` around `process_ubatch()` without requesting a scheduler reserve. [source]
- PR 28927 lists as untested: audio chunks, M-RoPE vision models such as Qwen VL, CPU-only backends and `--no-mmproj-offload`. [source]
- Tag b11232 was released on 2026-09-28 and is built from commit 6f767fe ("ggml-cpu: enable tiled flash attention for non-vector-multiple head dims on x86"). [source]
- A GitHub compare from ed7ac35 to b11232 lists 5 commits and 15 changed files, which means b11232 is ahead of the PR 28751 merge commit. [source]
- A GitHub compare from ed7ac35 to b11081 reports nothing to compare because ed7ac35 already contains all commits from b11081, so b11081 predates the merge. [source]
- Ollama tags v0.35.0 (pin b11081) lack the fix; v0.35.1 (pin b11232, see llama-cpp-bundled-revision-inside-ollama-release.md) includes it. [source]
- A Mac user sending several images per request to an Ollama build older than v0.35.1 pays two scheduler reserves per image; the cost is larger with a large `num_ctx`. [source]
Children
- No children recorded.