Mac local LLMs: Prompt cache and persistent KV
Parent: Running LLM models locally on a Mac · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Claude Code attribution block `x-anthropic-billing-header: cc_version=<ver>.<hash>; cc_entrypoint=cli; cch=00000;` is the first system block; in v2.1.37 the hash changed per call. Set `CLAUDE_CODE_ATTRIBUTION_HEADER=0` (settings.json env or `--settings`; older releases ignored shell export). Repo...
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Agent client invalidators (fix first)
- Claude Code attribution block `x-anthropic-billing-header: cc_version=<ver>.<hash>; cc_entrypoint=cli; cch=00000;` is the first system block; in v2.1.37 the hash changed per call. Set `CLAUDE_CODE_ATTRIBUTION_HEADER=0` (settings.json env or `--settings`; older releases ignored shell export). Reported 2 s follow-ups became 30 s waits. [source]
- Since v2.1.181 the block is stable within one conversation, but still differs across conversations and subagents, so cross-conversation system-prompt reuse still breaks. [source]
- Bedrock rejects the block: `400 x-anthropic-billing-header is a reserved keyword and may not be used in the system prompt`. [source]
- `--tools "Bash,Edit,Read"` plus `--strict-mcp-config` cut a 259-tool request to 8 built-ins; `--exclude-dynamic-system-prompt-sections` moves per-user context to the first user message (ignored with --system-prompt). [source]
- Thinking: Qwen3.5 template emits `<think>` only after the last user message, so history re-renders when a new user turn arrives. Qwen3.6 needs `--chat-template-kwargs '{"preserve_thinking":true}'`. [source]
- Non-leading system messages: oMLX `preserve_mid_system_cache` defaults true (user_note_safe); false merges all to one leading system message. [source]
- Template patching: llama.cpp `--jinja --chat-template-file chat_template.jinja --reasoning-format deepseek`; check with the props endpoint `chat_template`; froggeric's v22.5 ships render-twice-and-diff scripts. Diff the rendered tokenised prompt, not request JSON. [source]
llama.cpp / llama-server
- Defaults: `--ctx-checkpoints 32` per slot, `--checkpoint-min-step 8192`, `--cache-idle-slots` on, `--kv-unified` on with auto slots, `--slot-prompt-similarity 0.10`. Cache size via `--cache-ram`/`-cram`; log line `cache state: N prompts, X MiB (limits: ...)`. [source]
- Debug: `-lv 4` and grep `forcing full prompt|context checkpoint|Checking checkpoint|n_past =`; good `restored context checkpoint`, bad `erased invalidated context checkpoint`. No mismatch dump when `n_past == 0` (first-token change looks cold). [source]
- A 121,035-token prompt was reprocessed repeatedly once the cache hit its 8192 MiB limit; an evicted entry takes its checkpoints with it. Cache restore does not depend on slot id. [source]
- Checkpoint RAM: Gemma 4 on build 8724 went 0.7 GB to 10 GB to OOM above 25 GB; `--ctx-checkpoints 1 -np 1` held 1.5 GB, `0` held 0.4 GB. 128 checkpoints at 149.626 MiB is about 18.7 GiB per slot. [source]
- Restore search walks newest to oldest; if every checkpoint has `pos_min` above the threshold the whole prompt reprocesses (issue 23013). Upstream eviction is oldest-first plus min-step proximity (PR 25472), not tokens-per-byte scoring. [source]
- Recurrent models: `--cache-reuse` and `--swa-full` do nothing for Qwen3-Next-style; Gemma-3 without `--swa-full` restores "successfully" but re-prefills. With `--swa-full` a 21K prefix fell 230.7 s cold to 6.9 s warm. [source]
- Sleep (`--sleep-idle-seconds`): `prompt_cache` stays allocated while asleep but wake runs `load_model()`, which rebuilds it and drops saved states (issue 29322; PR 29408 adds `test_server_sleep_prompt_cache`). [source]
Persistence across restart, swap, unload
- Slot save files hold token list plus KV cells, not checkpoints; restore calls `slot->prompt.clear()`, so a restored hybrid slot shows `cache_n = 0` (issue 25913). PR 26004 appends an `SCKP` checkpoint payload (old files still load): 58,202 tokens 181.9 s to 4.7 s on Vulkan. Unmerged as of last read. [source]
- Save size ~128 KiB/token f16 (Llama-3.1-8B); q8_0 is 0.531x, q4_0 0.281x. [source]
- stillwarm refuses restore on flash-attn change, KV type change, model change, `--swa-full` change, or context smaller than saved tokens; `-ngl` changes are allowed; it warns when a restore succeeds but the server re-prefills. [source]
- llama-save-wrapper: `main.py --port 12872 -m model.gguf --slot-save-path DIR -c 32768`. llama-swap has no per-model hooks (discussion 724); a restore script races the first request because readiness is the health check. `/api/models/unload[/<model>]`, `unloadTimeout` default 10 s. [source]
- LM Studio disk cache is temporary `/tmp` overflow, not a restart tier. 0.3.35 broke MLX caching (revert to 0.3.31); 0.4.2 logs `Tried to trim '3195' tokens from the prompt cache, but could not: Cache is not trimmable. Clearing the cache instead.` [source]
MLX servers with SSD tiers
- oMLX: SSD limit `auto` = 50% of (free disk + existing cache incl. GDN sidecars) (v0.6.4 used 10% of capacity); fixed `--paged-ssd-cache-max-size 20GB`. M4 Pro, 22,819 tokens: cold 63.1 s, later turns 0.49 s, after restart 0.27 s. [source]
- oMLX memory guard: `Prefill would require ~9.08 GB peak ... dynamic ceiling is 9.00 GB` on 18 GB. No per-hour write cap exists; wear is unmeasured. [source]
- Rapid-MLX: `--kv-disk-checkpoint-interval` is write-only (external tooling) and blocks decode O(context); `--hybrid-cache-entries N` default 0. Restart restored 18,944 of 22,819 tokens. Shutdown save has a 3.5 s SIGTERM budget. [source]
- MTPLX: `--ssd-session-cache-max-size` default 100GB (32GB on 64 GB Macs), writes stop below 10 GiB free; reasoning history preserved by default. [source]
- ds4: files keyed by SHA1 byte prefix; trims 32 tokens then rounds down to a multiple of 2,048; eviction score (hits with 6 h half-life + 1) x tokens / size; `--tool-memory-max-ids`, `--disable-exact-dsml-tool-replay`. [source]
Decision rules
- Agent slow on every turn: fix the client prefix first, then engine. Restarts frequent or multi-agent: oMLX or MTPLX SSD tier. Rare restarts: RAM cache (`--parallel 1` contends between chats). [source]
Corrections
Open
- Whether vLLM hybrid/Mamba prefix reuse is default (docs say WIP); CachyLlama design unreadable; oMLX on-disk size vs f16 unmeasured. [source]
Corrections and disagreements
- Does LM Studio persist KV across restarts? rapidmlx.com comparison table lists LM Studio as "In-memory". The LM Studio blog describes a disk-backed cache that is intentionally temporary (`/tmp` scratch file, cleared on model unload). Both are consistent: disk is used as overflow inside a model lifetime, not as a restart tier. CONTRADICTS prompt-processing-versus-decode-on-apple-gpus-and-neural-accelerators.md line 23 ("appear to survive restarts and evictions" for LM Studio mlx-engine 1.8.5); the survivors are oMLX and, partially, Rapid-MLX. [source]
- CONTRADICTS prompt-processing-versus-decode-on-apple-gpus-and-neural-accelerators.md line 23: mlx-engine 1.8.5's disk cache does not survive restarts. [source]
- Gemma 4 26B-A4B depth (extends round 1): the llama.cpp load log shows `block_count 30` with a 4,608-cell SWA cache, matching the mlx discussion (30 layers) and CONTRADICTS the "10 global / 50 sliding" claim, which implies 60 layers. [source]
- Is the attribution block per-request or per-conversation? Unsloth (and Rallo, Mykola) describe a value that changes every request and cost "90% slower", "2-second follow-ups into 30-second waits". Claude Code docs say stable per conversation since v2.1.181 and per-request before. Reporters do not state their version. CONTRADICTS: local-llm-server-as-a-coding-agent-backend-on-mac.md which says the line changes every request without a version bound. [source]
- CONTRADICTS: https://danmackinlay.name/notebook/local_llm_mac.html (calls `--kv-disk-checkpoint-interval` the nearest replacement for oMLX's persistent cache; not in an existing dossier): Rapid-MLX's CLI reference says the flag is write-only today and 'enable only for external tooling that consumes the files'. [source]
- CONTRADICTS: prompt-cache-survival-across-model-unload-swap-and-restart.md line "sleep unloads the model and its KV cache" implies host cache is freed; in master the `prompt_cache` object stays allocated during sleep and is discarded only on wake. [source]
- CONTRADICTS: hybrid-and-sliding-window-attention-kv-cache-rewinding.md (line 109, vLLM 0.28.0 enabled hybrid/Mamba prefix reuse by default): the current vLLM design page still says prefix caching for Mamba models is work in progress; the page may be stale. [source]
- CONTRADICTS: chat-template-patching-for-cache-stability.md (open question "no source documents a render-twice-and-diff procedure"). froggeric's repo ships exactly that, in two scripts. [source]
- Premise check, CONTRADICTS the concept name: no upstream file read here scores checkpoints by tokens per byte or decays a hit count. Upstream uses oldest-first plus min-step proximity. If such a scorer exists it is in a fork (the batch brief links it to CachyLlama), and no readable source documents it. [source]
Concepts in this cluster
- Disk and SSD-tiered KV cache servers [source]
- Hybrid and sliding-window attention KV cache rewinding [source]
- Prompt-cache invalidation by agent clients [source]
- Chat-template patching for cache stability [source]
- Context shift and cache-overflow policy for recurrent models [source]
- Prompt cache survival across model unload swap and server restart [source]
- Rapid-MLX server and kv-disk-checkpoint-interval [source]
- Rapid-MLX shutdown-time disk save [source]
- Reasoning-token cache invalidation across turns [source]
- llama-server host-memory prompt cache and idle sleep unload [source]
- stillwarm llama-server KV persistence wrapper [source]
- vLLM-Metal hybrid KV cache manager and prefix caching [source]
- Agent-owned KV cache export import artifacts [source]
- ds4 disk KV cache and SHA1 byte-prefix keying [source]
- llama-swap lifecycle hooks and wrapper-based KV save restore [source]
- llama.cpp host-memory prompt cache cache-ram cache-idle-slots on unified memory [source]
- llama.cpp host prompt cache interaction with id_slot pinning [source]
- Prefix-cache-safe handling of non-leading system turns [source]
- Prefix-stability regression testing of chat templates [source]
- CachyLlama llama.cpp fork with persistent KV [source]
- Exact tool-call replay for byte-stable agent prefixes [source]
- KV checkpoint eviction scoring by tokens per byte with hit decay [source]
- LLAMA_SERVER_SLOTS_DEBUG mismatch logging workflow [source]
- MTPLX SSD session cache and SessionBank warm-prefix reuse [source]
- llama.cpp prompt-cache eviction when cache-ram limit is hit [source]
- Chat-template and tool-declaration determinism as a cache-hit prerequisite [source]
- SSD write budgets and wear for KV session stores [source]
- oMLX paged SSD cache auto limit policy [source]
Children
- MTPLX SSD session cache and SessionBank warm-prefix reuse
- oMLX paged SSD cache auto limit policy
- Prefix-cache-safe handling of non-leading system turns
- Prefix-stability regression testing of chat templates
- Prompt-cache invalidation by agent clients
- Prompt cache survival across model unload swap and server restart
- Rapid-MLX server and kv-disk-checkpoint-interval
- Rapid-MLX shutdown-time disk save
- Reasoning-token cache invalidation across turns
- SSD write budgets and wear for KV session stores
- stillwarm llama-server KV persistence wrapper
- vLLM-Metal hybrid KV cache manager and prefix caching
- Agent-owned KV cache export import artifacts
- CachyLlama llama.cpp fork with persistent KV
- Chat-template and tool-declaration determinism as a cache-hit prerequisite
- Chat-template patching for cache stability
- Context shift and cache-overflow policy for recurrent models
- Disk and SSD-tiered KV cache servers
- ds4 disk KV cache and SHA1 byte-prefix keying
- Exact tool-call replay for byte-stable agent prefixes
- Hybrid and sliding-window attention KV cache rewinding
- KV checkpoint eviction scoring by tokens per byte with hit decay
- llama.cpp host-memory prompt cache cache-ram cache-idle-slots on unified memory
- llama.cpp host prompt cache interaction with id_slot pinning
- llama.cpp prompt-cache eviction when cache-ram limit is hit
- llama-server host-memory prompt cache and idle sleep unload
- LLAMA_SERVER_SLOTS_DEBUG mismatch logging workflow
- llama-swap lifecycle hooks and wrapper-based KV save restore