Context shift and cache-overflow policy for recurrent models
Parent: Mac local LLMs: Prompt cache and persistent KV · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
llama.cpp classifies a context by `llama_memory_seq_rm`: partial removal works (no checkpoint needed), whole-sequence only (checkpoint needed), or bounded rollback for recurrent state (checkpoint needed past the bound).
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- llama.cpp classifies a context by `llama_memory_seq_rm`: partial removal works (no checkpoint needed), whole-sequence only (checkpoint needed), or bounded rollback for recurrent state (checkpoint needed past the bound). [source]
- Recurrent layers keep one running state overwritten each step; a short ring of per-token snapshots allows rollback only within `n_rs_seq` tokens. [source]
- Hence shifting or dropping the head of a prompt cannot be expressed for a hybrid model; `--context-shift` is off by default and `--cache-reuse` is disabled for such contexts. [source]
- 2023-11 (issue 3969, attention models): a context-shift loop when `n_keep` left almost no room; reporters proposed capping discards. [source]
- 2026: hybrid models make the overflow policy a client concern (compaction), not a server feature. [source]
- A request larger than the slot context is rejected and the conversation can lose its slot, evicting its own warm state. [source]
- When context fills and the cache limit forces eviction, a hybrid slot reprocesses the whole history. [source]
- Speculative decoding with partial acceptance can need a rollback beyond `n_rs_seq`; PR 24797 makes that fall through to normal cell removal rather than fail. [source]
- Operators disable shifting (`--no-context-shift`) for correctness with hybrid models; some Qwen3.6 reporters kept it on and never reached a diagnosis (23030, held elsewhere). [source]
- Whether any llama.cpp change plans a recurrent-aware shift; none was found in the sources. [source]
- The server README lists `--context-shift` as "whether to use context shift on infinite text generation" with default disabled (env `LLAMA_ARG_CONTEXT_SHIFT`). [source]
- Particula's source reading classifies contexts as `SEQ_RM_TYPE_NO`, `PART` (arbitrary partial rollback, no checkpoint), `FULL` (whole sequences only, checkpoint needed) and `RS` (partial rollback bounded by `n_rs_seq`, checkpoint needed past the bound). [source]
- The recurrent rewind succeeds only when `cell.pos - (p0 - 1)` lies between 1 and `n_rs_seq`, via a ring of per-token snapshots. [source]
- Checkpoint creation is gated on the context being `FULL` or `RS`, or the model having `n_swa > 0`. [source]
- A Qwen3.6-27B hybrid load logs `cache_reuse is not supported by this context, it will be disabled` and `swa_full is not supported by this model, it will be disabled`. [source]
- That same load logs `speculative decoding will use checkpoints` for the MTP-enabled hybrid model. [source]
- A working 24055 server line for Qwen3.6-27B uses `--no-context-shift` with `--parallel 1`. [source]
- Particula advises compacting before the slot limit because an oversized request that gets rejected can leave the conversation without its slot, evicting its warm state. [source]
- PR 24797 changes hybrid `seq_rm`: when speculative decoding accepts part of a draft and the rollback distance exceeds `n_rs_seq`, the recurrent cache now falls through to normal cell removal instead of returning false, so attention KV is trimmed correctly but the recurrent state loses exact snapshot restoration. [source]
- Issue 3969 (b1492, 2023, deepseek-coder 33B, `-c 6120`) logged repeated `context shift - n_keep = 2224, n_left = 3894, n_discard = 1947` without progress; a reporter's patch capped `n_discard = std::min(n_left, 32)`. [source]
- Inferred: for hybrid models on a Mac, the overflow policy is client-side compaction before the context limit plus `--no-context-shift`; the server cannot drop the prompt head without a full re-prefill. [source]
Children
- No children recorded.