CachyLlama llama.cpp fork with persistent KV
Parent: Mac local LLMs: Prompt cache and persistent KV · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
The only primary page, the Reddit thread, is refused by the fetch helper ("we do not support this site"), so the fork's mechanism is not known from any source.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- The only primary page, the Reddit thread, is refused by the fetch helper ("we do not support this site"), so the fork's mechanism is not known from any source. [source]
- Upstream llama.cpp already persists KV through slot save/restore files, and a 2026-08-12 report shows a slot file restore bringing back context checkpoints: `load_slot_ch: restored 3 context checkpoint(s) from '.../4.bin'`, then a normal checkpoint restore at n_tokens 18347. [source]
- That report ran PR 25592 together with PR 26004 (the slot-file change that carries checkpoints), so persistence of checkpoints across a restart is an upstream-PR capability, not fork-only. [source]
- Inferred: a fork that advertises "persistent KV" most likely automates slot save/restore plus checkpoint persistence, the same ground as stillwarm and PR 26004. [source]
- The Reddit post id 1v5k08a places it after the 2026-07 stillwarm release; the date is not readable. [source]
- Other forks that carry hybrid-cache fixes ahead of upstream are named in the PR threads: BeeLlama (Anbeeld/beellama.cpp, recurrent shrink/expand), tultr/llama.cpp (applies PR 25592, publishes a CUDA image) and ik_llama.cpp. [source]
- Fork-specific features cannot be verified against upstream source, so claims about CachyLlama's eviction scoring, slot trimming or exact tool-call replay (the other concepts in this batch) stay unverified here. [source]
- A fork that tracks upstream loses fixes when upstream refactors prompt-cache ownership, as PR 25649 did on 2026-07-14 by moving KV state into `server_prompt_cache_state`. [source]
- None on record: no second source discusses CachyLlama. [source]
- What CachyLlama persists (slot files, checkpoints, host cache), where, and under which cache key. No readable source says. [source]
- Whether the checkpoint eviction score (tokens per byte with hit decay), user-boundary spans, reasoning-trim and exact tool-call replay in this batch are CachyLlama features. The batch brief groups them; no source confirms it. [source]
- A second pass needs a reader that can fetch Reddit, or the fork's GitHub README. [source]
- The fetch helper refuses the CachyLlama Reddit thread, so its design is unreadable in this corpus. [source]
- Search for "CachyLlama github" and for its design terms returned only the Reddit URL and unrelated pages, so no repository page was found. [source]
- A 2026-08-12 PR 25592 comment shows slot-file restore returning 3 context checkpoints (`load_slot_ch: restored 3 context checkpoint(s)`) on a hybrid Qwen3.6-27B, run with PRs 25592 and 26004. [source]
- The same reporter says model switches and resumed work no longer wait the roughly 75 s that 60k tokens of prefill takes at 800 tokens per second. [source]
- BeeLlama (Anbeeld/beellama.cpp) is the origin of the recurrent_shrink/recurrent_expand mechanism that PR 24785 tried to backport. [source]
- tultr/llama.cpp merged a PR applying upstream PR 25592 and publishes `ghcr.io/tultr/llama.cpp:master`. [source]
- PR 25649 moved prompt KV data from `server_prompt` into a `server_prompt_cache_state` owned by `server_prompt_cache` and merged on 2026-07-14. [source]
Children
- No children recorded.