llama.cpp host-memory prompt cache cache-ram cache-idle-slots on unified memory
Parent: Mac local LLMs: Prompt cache and persistent KV · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
The cache has a byte limit and a token limit (default: context size); either can evict an entry.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- The cache has a byte limit and a token limit (default: context size); either can evict an entry. [source]
- Each hybrid entry includes recurrent state plus checkpoints, so one entry can be hundreds of MiB. [source]
- When an entry is evicted, a later request for that prompt restores nothing and, on a hybrid model, also has no checkpoints. [source]
- Eviction cascade: a user on a 262144-context hybrid run with `--cache-ram 8192` saw a 121,035-token reprocess each time a cached prompt was dropped. [source]
- Over-asking checkpoints: a compose file with `--ctx-checkpoints 64 --cache-ram 2048` would need about 9.35 GiB of 149.626 MiB checkpoints per slot against a 2 GiB cache. [source]
- No Mac measurement of a good `--cache-ram` per RAM tier exists in the sources. [source]
- Whether checkpoints count against `--cache-ram` while still attached to a live slot is not stated in the cited pages. [source]
- PR 16391 gives `-cram -1` as no limit ("as much host RAM as is available") and `-cram 0` as off, with a second limit on cached tokens that defaults to the context size. [source]
- A b9354 log shows the effective limits on every cache update: `limits: 2048.000 MiB, 32768 tokens, 2147483648 est`. [source]
- A 24055 commenter ties repeated 121,035-token reprocessing to the prompt cache reaching its 8192 MiB limit (about 7.9 GiB used) and dropping a cached prompt. [source]
- Observed checkpoint sizes of 149.626, 175.502 and 254.797 MiB for 27B hybrid variants make a default 32-checkpoint slot 4.7 to 8 GiB, comparable to the 8192 MiB default cache. [source]
- The 24055 reproducer set `--ctx-checkpoints 64` with `--cache-ram 2048` and 149.626 MiB checkpoints; 64 of them total about 9.35 GiB. [source]
- Inferred: on a 16 to 24 GB Mac, the 8192 MiB default takes a third to half of RAM, so use `-cram 0` while measuring and a deliberate value afterward; on 64 GB and above leave the default and size `-ctxcp` and `-cms` so checkpoints fit inside it. [source]
- Inferred: because `--cache-idle-slots` clears an idle slot into the cache under unified KV, a single-slot Mac setup (`-np 1`) gets an LRU-like cache of past conversations only if `--cache-ram` is large enough for at least two full states. [source]
Children
- No children recorded.