<!-- llms-explorer concept facts · https://llms-explorer.com/tree/cachyllama-llama-cpp-fork-with-persistent-kv/ · pack 2026-10-05 · ~1092 tokens -->

# CachyLlama llama.cpp fork with persistent KV

> The only primary page, the Reddit thread, is refused by the fetch helper ("we do not support this site"), so the fork's mechanism is not known from any source.

Parent: [Mac local LLMs: Prompt cache and persistent KV](https://llms-explorer.com/tree/mac-local-llms-prompt-cache-and-persistent-kv/) · 1 facets · 19 facts · page: https://llms-explorer.com/tree/cachyllama-llama-cpp-fork-with-persistent-kv/

## Facts

- The only primary page, the Reddit thread, is refused by the fetch helper ("we do not support this site"), so the fork's mechanism is not known from any source. — [source](https://www.reddit.com/r/LocalLLaMA/comments/1v5k08a/cachyllamas_llamacpp_fork_with_persistent_kv/)
- Upstream llama.cpp already persists KV through slot save/restore files, and a 2026-08-12 report shows a slot file restore bringing back context checkpoints: `load_slot_ch: restored 3 context checkpoint(s) from '.../4.bin'`, then a normal checkpoint restore at n_tokens 18347. — [source](https://github.com/ggml-org/llama.cpp/pull/25592)
- That report ran PR 25592 together with PR 26004 (the slot-file change that carries checkpoints), so persistence of checkpoints across a restart is an upstream-PR capability, not fork-only. — [source](https://github.com/ggml-org/llama.cpp/pull/25592)
- Inferred: a fork that advertises "persistent KV" most likely automates slot save/restore plus checkpoint persistence, the same ground as stillwarm and PR 26004. — source: `asserted`
- The Reddit post id 1v5k08a places it after the 2026-07 stillwarm release; the date is not readable. — source: `asserted`
- Other forks that carry hybrid-cache fixes ahead of upstream are named in the PR threads: BeeLlama (Anbeeld/beellama.cpp, recurrent shrink/expand), tultr/llama.cpp (applies PR 25592, publishes a CUDA image) and ik_llama.cpp. — [source](https://github.com/ggml-org/llama.cpp/pull/24785)
- Fork-specific features cannot be verified against upstream source, so claims about CachyLlama's eviction scoring, slot trimming or exact tool-call replay (the other concepts in this batch) stay unverified here. — source: `asserted`
- A fork that tracks upstream loses fixes when upstream refactors prompt-cache ownership, as PR 25649 did on 2026-07-14 by moving KV state into `server_prompt_cache_state`. — [source](https://github.com/ggml-org/llama.cpp/pull/25649)
- None on record: no second source discusses CachyLlama. — source: `asserted`
- What CachyLlama persists (slot files, checkpoints, host cache), where, and under which cache key. No readable source says. — source: `asserted`
- Whether the checkpoint eviction score (tokens per byte with hit decay), user-boundary spans, reasoning-trim and exact tool-call replay in this batch are CachyLlama features. The batch brief groups them; no source confirms it. — source: `asserted`
- A second pass needs a reader that can fetch Reddit, or the fork's GitHub README. — source: `asserted`
- The fetch helper refuses the CachyLlama Reddit thread, so its design is unreadable in this corpus. — [source](https://www.reddit.com/r/LocalLLaMA/comments/1v5k08a/cachyllamas_llamacpp_fork_with_persistent_kv/)
- Search for "CachyLlama github" and for its design terms returned only the Reddit URL and unrelated pages, so no repository page was found. — source: `asserted`
- A 2026-08-12 PR 25592 comment shows slot-file restore returning 3 context checkpoints (`load_slot_ch: restored 3 context checkpoint(s)`) on a hybrid Qwen3.6-27B, run with PRs 25592 and 26004. — [source](https://github.com/ggml-org/llama.cpp/pull/25592)
- The same reporter says model switches and resumed work no longer wait the roughly 75 s that 60k tokens of prefill takes at 800 tokens per second. — [source](https://github.com/ggml-org/llama.cpp/pull/25592)
- BeeLlama (Anbeeld/beellama.cpp) is the origin of the recurrent_shrink/recurrent_expand mechanism that PR 24785 tried to backport. — [source](https://github.com/ggml-org/llama.cpp/pull/24785)
- tultr/llama.cpp merged a PR applying upstream PR 25592 and publishes `ghcr.io/tultr/llama.cpp:master`. — [source](https://github.com/ggml-org/llama.cpp/pull/25592)
- PR 25649 moved prompt KV data from `server_prompt` into a `server_prompt_cache_state` owned by `server_prompt_cache` and merged on 2026-07-14. — [source](https://github.com/ggml-org/llama.cpp/pull/25649)
