llama.cpp host prompt cache interaction with id_slot pinning
Parent: Mac local LLMs: Prompt cache and persistent KV · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Unpinned requests go through slot selection (`get_availabl`): LRU or LCP-similarity choice, then a prompt-cache update that looks for a better cached prompt and swaps it in.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Unpinned requests go through slot selection (`get_availabl`): LRU or LCP-similarity choice, then a prompt-cache update that looks for a better cached prompt and swaps it in. [source]
- Pinned requests name the slot up front; the known result (26004) is that the cache swap-in did not help them, consistent with the selection step being skipped. [source]
- Cached prompts are not tied to slot ids: a prompt can come back in a different slot and still be restored. [source]
- With `-np 1` there is only slot 0, so pinning adds nothing and the host cache supplies all reuse. [source]
- With `--cache-idle-slots` on and unified KV, an idle slot is saved to the cache and cleared when a new task starts, so a pinned slot may hold nothing when it is next addressed. [source]
- No source states in code or docs whether `id_slot` bypasses the cache load by design or by omission. [source]
- The server README defines `id_slot` as assigning the completion task to a specific slot, with -1 (the default) meaning any idle slot. [source]
- The server README says `--cache-idle-slots` saves idle slots to the prompt cache on a new task and clears them when using unified KV. [source]
- In a b9354 log the slot pick and the cache lookup are consecutive steps: `selected slot by LRU, t_last = -1`, then `get_availabl: updating prompt cache`, then `looking for better prompt, base f_keep = -1.000, sim = 0.000`. [source]
- When a prompt shares a prefix, the pick is `selected slot by LCP similarity, sim_best = 0.928 (> 0.100 thold), f_keep = 0.911`. [source]
- In the 29322 four-slot log, two prompts first processed in slots 2 and 1 came back later in slots 0 and 3 with 4 tokens processed, so cache restore does not depend on the original slot id. [source]
- PR 16391 describes the host cache as "extra slots" whose prefix similarity is computed against the incoming prompt and which are hot-swapped into the context when that saves work. [source]
- Inferred: a harness that pins `id_slot` for its own save/restore should either also keep `-np` at the number of concurrent sessions or leave `id_slot` unset and let `--cache-ram` do prefix reuse. [source]
Children
- No children recorded.