llama.cpp seq_pos_min hybrid memory fix PR 24797
Parent: Mac local LLMs: llama.cpp internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Master still reads `seq_pos_min` as `std::max(mem_attn->seq_pos_min(seq_id), mem_recr->seq_pos_min(seq_id))` with the comment "the min of the total cache is the max of the two caches' min values", and `seq_pos_max` as the `std::min` of both.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Master still reads `seq_pos_min` as `std::max(mem_attn->seq_pos_min(seq_id), mem_recr->seq_pos_min(seq_id))` with the comment "the min of the total cache is the max of the two caches' min values", and `seq_pos_max` as the `std::min` of both. [source]
- The reason ggerganov gave for rejecting the change: "the recurrent state is valid only for the latest position and using it with any other position is incorrect". [source]
- The rejected change also let a rollback that exceeded `n_rs_seq` fall through to normal cell removal, which keeps attention KV trimmed but leaves a recurrent state at the wrong position. [source]
- The replacement PR 25592 states the invariant instead: recurrent state is only restored at the exact position it was saved at. It records checkpoint validity at save time (`pos_min = pos_max` for hybrid and recurrent memory), restores a checkpoint only when its exact position lies inside the common prefix of the new prompt, and erases checkpoints that cover diverged content instead of trimming them. [source]
- PR 25592 makes no change to `llama_memory_hybrid`; `seq_rm` still fails cleanly on impossible rollbacks, now with a debug log. [source]
- The master server code carries the open marker: `[TAG_CHECKPOINTS_FIX_POS_MIN]` TODO "here we incorrectly determine that the saved checkpoint data covers the [pos_min, pos_max] range", plus a restore-side workaround that rejects a checkpoint with `pos_max > pos_next`. [source]
- PR 24110 changed `pos_min_thold` so the `-1` margin applies only when no new tokens exist (`has_new_tokens = n_past < task.n_tokens()`); this avoids restoring a checkpoint when the slot already holds the prefix. Master has it. [source]
- 2026-06-19: PR 24797 opened. 2026-06-26/27: a tester reports most invalidations gone on b9820 with the patch, still seeing `clearing stale recurrent state beyond n_past = 10047 (recr_pos_max = 15932)` under many parallel sessions. [source]
- 2026-06-27: ggerganov calls the change "obviously wrong" and closes the PR. [source]
- 2026-07-12: krim404 opens PR 25592 as "a reimplementation of #24797" and states that 24797 "was rejected because it could reuse recurrent state at positions it was never valid for". [source]
- 2026-07-14: the PR is rebased after PRs 25472 and 25649 landed; a second commit drops the bounded `n_rs_seq` rollback fast path because those snapshots are valid only for tokens decoded in the last ubatch. [source]
- 2026-07 to 2026-08: testers report 150 to 400 times faster turns on restored checkpoints; the newest comment on the cached page is dated 2026-08-26 and no merge event appears. [source]
- Speculative decoding can loop on hybrid models: PR 25819 (WIP mitigation, 2026-07-17) logs `STUCK speculative loop: 4 consecutive checkpoint restores with no progress` for ngram-mod drafts after a restore. A 2026-08-25 commenter says the bug "happens often enough to be noticeable" with 25592 and 26004. [source]
- PR 24785, a different approach (recurrent shrink/expand ported from BeeLlama), was also proposed for the same symptoms. [source]
- Contributor view: the max() inflates `pos_min`, so the fix is to report the attention minimum. Maintainer view: positions other than the latest are invalid for recurrent state, so the fix belongs in checkpoint metadata and restore rules. They disagree on where the invariant is enforced; the maintainer view decided the outcome. [source]
- Testers of 24797 reported working caches; the closure reason says the cache hits could rest on invalid state. Whether any corruption occurred is unmeasured; 25592 reports byte-identical output at temperature 0 against a cold prefill. [source]
- Whether PR 25592 merges, and whether upstream adopts its `pos_min = pos_max` bookkeeping or a different one. [source]
- Whether Metal behaves like the CUDA, ROCm and Vulkan logs; every cited test is on those backends. [source]
- PR 24797 was closed by ggerganov on 2026-06-27 and is not merged. [source]
- ggerganov's closing comment is "This change is obviously wrong - the recurrent state is valid only for the latest position and using it with any other position is incorrect." [source]
- Master `llama_memory_hybrid::seq_pos_min` still returns the max of the attention and recurrent minimums. [source]
- PR 25592 describes itself as a reimplementation of 24797 that keeps the exact-position invariant. [source]
- PR 25592 resolves the `[TAG_CHECKPOINTS_FIX_POS_MIN]` TODO for hybrid and recurrent memory. [source]
- PR 25592 logs checkpoint create, restore and erase at INFO so they show without `-lv`. [source]
- PR 25592's author tested on Qwen3.6-35B-A3B Q4_K_M on Vulkan; a repeated 22.8k-token request went from about 96 s to a 4-token cache hit. [source]
- A tester on an RTX 3090 with Qwen3.8-27B IQ4_NL measured a 38.5K-token prefill of 40.8 s falling to 0.1 to 0.3 s after restore. [source]
- A third-party report says reusing an equivalent checkpoint instead of saving a new one cut a 1156-token warm request from about 122 ms to 43 ms on a Radeon 780M. [source]
- Issue 22746 (Qwen 3.6 27B full re-processing, opened 2026-05-06) is marked Closed. [source]
- Issue 23589 traced a cache drop of one batch (`-b`) worth of tokens per turn to commit ccee426 (PR 23280). [source]
- The author of PR 24110 said it might also fix issue 23589, and ggerganov called the change technically correct but asked for a repro on master. [source]
Corrections and disagreements
- CONTRADICTS: llama-cpp-context-checkpoints-and-swa-hybrid-pro.md ("2026-06-19: PR 24797 opened to fix seq_pos_min(); draft/open at fetch") and its open question "Whether 24797 merges": it was closed unmerged on 2026-06-27. [source]
- CONTRADICTS: hybrid-gateddeltanet-state-rollback-and-checkpoi.md (PR 24797 "draft", "targets the `seq_pos_min()` root cause") and its claims that the `std::max()` is a bug to fix; the maintainer treats the recurrent state's single valid position as the invariant. [source]
- CONTRADICTS: context-shift-and-cache-overflow-policy-for-recu.md ("PR 24797 makes that fall through to normal cell removal"): that behavior never merged. [source]
Children
- No children recorded.