<!-- llms-explorer concept facts · https://llms-explorer.com/tree/llama-cpp-pr-25592-exact-position-hybrid-checkpo/ · pack 2026-10-05 · ~1285 tokens -->

# llama.cpp PR 25592 exact-position hybrid checkpoint restore

> The PR page shows state Open, label `server`, "Awaiting requested review from ggerganov" and "At least 2 approving reviews are required to merge this pull request".

Parent: [Mac local LLMs: llama.cpp internals](https://llms-explorer.com/tree/mac-local-llms-llama-cpp-internals/) · 2 facets · 18 facts · page: https://llms-explorer.com/tree/llama-cpp-pr-25592-exact-position-hybrid-checkpo/

## Facts

- The PR page shows state Open, label `server`, "Awaiting requested review from ggerganov" and "At least 2 approving reviews are required to merge this pull request". — [source](https://github.com/ggml-org/llama.cpp/pull/25592)
- On 2026-09-04 a user running a cherry-picked patch on build 10555 (commit ea4787555) with Qwen3.8-27B at 130K context measured a 37.3K-token prompt: full prefill about 35 s, then every later turn restored from a checkpoint in 1.3 s (about 27 times faster). — [source](https://github.com/ggml-org/llama.cpp/pull/25592)
- That report shows the restore logged at INFO with equal pos_min and pos_max (37273=37273), meaning the recurrent state is restored only at its exact saved position. — [source](https://github.com/ggml-org/llama.cpp/pull/25592)
- A 2026-08-23 test of the equivalent-checkpoint reuse part, on llama.cpp v0.2.0 (bb4caa7) with Qwen3.5-2B IQ4_NL on a Radeon 780M under Vulkan, cut a warm request (1156 cached, 4 evaluated tokens) from about 122 ms to about 43 ms by removing a redundant checkpoint save of about 70 ms. — [source](https://github.com/ggml-org/llama.cpp/pull/25592)
- A second reuse case in the same report went from 157.3 ms to 73.2 ms. — [source](https://github.com/ggml-org/llama.cpp/pull/25592)
- A commenter summarising a local deployment says the PR "sets pos_min = pos_max for hybrid/recurrent models" and adds "checkpoint adoption, restore, and erase at exact positions with INFO logging". — [source](https://github.com/ggml-org/llama.cpp/pull/25592)
- One commenter reports the PR together with PR 26004 makes checkpoints and slots "finally usable with Qwen3.6", and another reports that with both PRs the issue 25819 loop "happens often enough to be noticeable". — [source](https://github.com/ggml-org/llama.cpp/pull/25592)
- On 2026-08-26 a user running MTP only (`--spec-type draft-mtp --spec-draft-n-max 4`, no ngram-mod) with the patch saw no stuck-loop symptoms and suggested the loop is specific to ngram-mod. — [source](https://github.com/ggml-org/llama.cpp/pull/25592)
- PR 25819 is a work-in-progress mitigation: when ngram-mod speculative decoding fails verification, `spec_draft` is set to the accepted tokens and the checkpoint restored, and the next iteration reuses the draft and can fail again. — [source](https://github.com/ggml-org/llama.cpp/pull/25819)
- PR 25819 logs a `STUCK speculative loop` warning after 4 consecutive checkpoint restores with no progress (matched 20 of 21 draft tokens) and forces progress; its author says the root cause is not yet clear. — [source](https://github.com/ggml-org/llama.cpp/pull/25819)
- PR 26004 appends a tagged payload (magic `SCKP`, version, count, per-checkpoint position fields and state blobs) after the llama state payload in a slot save file, and `load_slot_checkpoints()` reattaches them on restore. — [source](https://github.com/ggml-org/llama.cpp/pull/26004)
- PR 26004 is backward and forward compatible (old files restore as before; old servers stop reading at the end of their payload), caps the checkpoint count at 1024, aborts cleanly on truncation, and changes `tools/server/server-context.cpp` by about 120 lines. — [source](https://github.com/ggml-org/llama.cpp/pull/26004)
- PR 26004 says it was reproduced by nine users on CUDA, Metal, Vulkan and an Ollama integration, including a cross-machine CUDA-to-Metal restore, and it fixes issue 25913. — [source](https://github.com/ggml-org/llama.cpp/pull/26004)
- PR 26004's test `test_slot_restore_preserves_context_checkpoints` fails on master with `assert 210 == 31` (the full prompt is reprocessed) and passes with the change. — [source](https://github.com/ggml-org/llama.cpp/pull/26004)
- Issue 24055 (opened 2026-06-03 on docker image server-cuda12-b9354, "Context checkpoints always invalidated on hybrid/recurrent models") is still open and is cross-referenced from PR 25592. — [source](https://github.com/ggml-org/llama.cpp/issues/24055)
- Because the PR is unmerged, the Ollama pin b11232 (2026-09-29) cannot contain it. — source: `asserted`
- Practical rule on a Mac with hybrid Qwen models: expect full re-prefill on prompt divergence from stock llama-server and stock Ollama, and use a build carrying PR 25592 (plus PR 26004 for model switches) when agent turns share long prefixes. — source: `asserted`

## Corrections and disagreements

- CONTRADICTS: llama-cpp-seq-pos-min-hybrid-memory-fix-pr-24797.md ("newest comment on the cached page is dated 2026-08-26 and no merge event appears"): the page fetched 2026-10-04 has comments through 2026-09-04 and fork references through 2026-09-23; it is still open. — [source](https://github.com/ggml-org/llama.cpp/pull/25592)
