<!-- llms-explorer concept facts · https://llms-explorer.com/tree/llama-cpp-host-memory-prompt-cache-cache-ram-cac/ · pack 2026-10-05 · ~750 tokens -->

# llama.cpp host-memory prompt cache cache-ram cache-idle-slots on unified memory

> The cache has a byte limit and a token limit (default: context size); either can evict an entry.

Parent: [Mac local LLMs: Prompt cache and persistent KV](https://llms-explorer.com/tree/mac-local-llms-prompt-cache-and-persistent-kv/) · 1 facets · 14 facts · page: https://llms-explorer.com/tree/llama-cpp-host-memory-prompt-cache-cache-ram-cac/

## Facts

- The cache has a byte limit and a token limit (default: context size); either can evict an entry. — source: `asserted`
- Each hybrid entry includes recurrent state plus checkpoints, so one entry can be hundreds of MiB. — source: `asserted`
- When an entry is evicted, a later request for that prompt restores nothing and, on a hybrid model, also has no checkpoints. — source: `asserted`
- Eviction cascade: a user on a 262144-context hybrid run with `--cache-ram 8192` saw a 121,035-token reprocess each time a cached prompt was dropped. — source: `asserted`
- Over-asking checkpoints: a compose file with `--ctx-checkpoints 64 --cache-ram 2048` would need about 9.35 GiB of 149.626 MiB checkpoints per slot against a 2 GiB cache. — source: `asserted`
- No Mac measurement of a good `--cache-ram` per RAM tier exists in the sources. — source: `asserted`
- Whether checkpoints count against `--cache-ram` while still attached to a live slot is not stated in the cited pages. — source: `asserted`
- PR 16391 gives `-cram -1` as no limit ("as much host RAM as is available") and `-cram 0` as off, with a second limit on cached tokens that defaults to the context size. — [source](https://github.com/ggml-org/llama.cpp/pull/16391)
- A b9354 log shows the effective limits on every cache update: `limits: 2048.000 MiB, 32768 tokens, 2147483648 est`. — [source](https://github.com/ggml-org/llama.cpp/issues/24055)
- A 24055 commenter ties repeated 121,035-token reprocessing to the prompt cache reaching its 8192 MiB limit (about 7.9 GiB used) and dropping a cached prompt. — [source](https://github.com/ggml-org/llama.cpp/issues/24055)
- Observed checkpoint sizes of 149.626, 175.502 and 254.797 MiB for 27B hybrid variants make a default 32-checkpoint slot 4.7 to 8 GiB, comparable to the 8192 MiB default cache. — [source](https://github.com/ggml-org/llama.cpp/issues/24055)
- The 24055 reproducer set `--ctx-checkpoints 64` with `--cache-ram 2048` and 149.626 MiB checkpoints; 64 of them total about 9.35 GiB. — source: `asserted`
- Inferred: on a 16 to 24 GB Mac, the 8192 MiB default takes a third to half of RAM, so use `-cram 0` while measuring and a deliberate value afterward; on 64 GB and above leave the default and size `-ctxcp` and `-cms` so checkpoints fit inside it. — source: `asserted`
- Inferred: because `--cache-idle-slots` clears an idle slot into the cache under unified KV, a single-slot Mac setup (`-np 1`) gets an LRU-like cache of past conversations only if `--cache-ram` is large enough for at least two full states. — source: `asserted`
