<!-- llms-explorer concept facts · https://llms-explorer.com/tree/llama-cpp-host-prompt-cache-interaction-with-id/ · pack 2026-10-05 · ~773 tokens -->

# llama.cpp host prompt cache interaction with id_slot pinning

> Unpinned requests go through slot selection (`get_availabl`): LRU or LCP-similarity choice, then a prompt-cache update that looks for a better cached prompt and swaps it in.

Parent: [Mac local LLMs: Prompt cache and persistent KV](https://llms-explorer.com/tree/mac-local-llms-prompt-cache-and-persistent-kv/) · 1 facets · 13 facts · page: https://llms-explorer.com/tree/llama-cpp-host-prompt-cache-interaction-with-id/

## Facts

- Unpinned requests go through slot selection (`get_availabl`): LRU or LCP-similarity choice, then a prompt-cache update that looks for a better cached prompt and swaps it in. — source: `asserted`
- Pinned requests name the slot up front; the known result (26004) is that the cache swap-in did not help them, consistent with the selection step being skipped. — source: `asserted`
- Cached prompts are not tied to slot ids: a prompt can come back in a different slot and still be restored. — source: `asserted`
- With `-np 1` there is only slot 0, so pinning adds nothing and the host cache supplies all reuse. — source: `asserted`
- With `--cache-idle-slots` on and unified KV, an idle slot is saved to the cache and cleared when a new task starts, so a pinned slot may hold nothing when it is next addressed. — source: `asserted`
- No source states in code or docs whether `id_slot` bypasses the cache load by design or by omission. — source: `asserted`
- The server README defines `id_slot` as assigning the completion task to a specific slot, with -1 (the default) meaning any idle slot. — [source](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)
- The server README says `--cache-idle-slots` saves idle slots to the prompt cache on a new task and clears them when using unified KV. — [source](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)
- In a b9354 log the slot pick and the cache lookup are consecutive steps: `selected slot by LRU, t_last = -1`, then `get_availabl: updating prompt cache`, then `looking for better prompt, base f_keep = -1.000, sim = 0.000`. — [source](https://github.com/ggml-org/llama.cpp/issues/24055)
- When a prompt shares a prefix, the pick is `selected slot by LCP similarity, sim_best = 0.928 (> 0.100 thold), f_keep = 0.911`. — [source](https://github.com/ggml-org/llama.cpp/issues/24055)
- In the 29322 four-slot log, two prompts first processed in slots 2 and 1 came back later in slots 0 and 3 with 4 tokens processed, so cache restore does not depend on the original slot id. — [source](https://github.com/ggml-org/llama.cpp/issues/29322)
- PR 16391 describes the host cache as "extra slots" whose prefix similarity is computed against the incoming prompt and which are hot-swapped into the context when that saves work. — [source](https://github.com/ggml-org/llama.cpp/pull/16391)
- Inferred: a harness that pins `id_slot` for its own save/restore should either also keep `-np` at the number of concurrent sessions or leave `id_slot` unset and let `--cache-ram` do prefix reuse. — source: `asserted`
