LLAMA_SERVER_SLOTS_DEBUG mismatch logging workflow
Parent: Mac local LLMs: Prompt cache and persistent KV · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
At load the server logs `LLAMA_SERVER_SLOTS_DEBUG = <n>` as a warning when it is non-zero, and `LLAMA_SERVER_SLOTS_N_DIFF = <n>` when that one is non-zero; seeing those two lines is the check that the environment reached the process.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- At load the server logs `LLAMA_SERVER_SLOTS_DEBUG = <n>` as a warning when it is non-zero, and `LLAMA_SERVER_SLOTS_N_DIFF = <n>` when that one is non-zero; seeing those two lines is the check that the environment reached the process. [source]
- The dump runs only when `n_past > 0 && n_past <= slot.prompt.n_tokens()`. If the new prompt shares no prefix with the slot (`n_past == 0`), nothing is printed. [source]
- Window: `np0 = max(n_past - N_DIFF, 0)` and `np1 = min(n_past + N_DIFF + 2, min(old size, new size))`. With N_DIFF unset (0) the dump shows only token indices `n_past` and `n_past + 1`. [source]
- Each token prints as its text piece (`common_token_to_piece`) and as an 8-wide id; a media placeholder prints as `[mtmd]`; a ` | ` marker is inserted at index `n_past`, the first differing position. [source]
- Four warnings follow in order: `old: ... <text>`, `new: ... <text>`, then the old token ids and the new token ids. [source]
- The dump comes before the checkpoint search, so it prints even when a checkpoint then restores successfully. [source]
- With the debug flag on, the slot list endpoint serializes with `only_metrics = false`, so each slot adds `prompt` (the detokenized task tokens) and `generated` (the generated text, or the copy saved at slot release). [source]
- At release the server copies `slot.generated_text` into `debug_generated_text` only when the flag is set. [source]
- ggerganov named the flag as the diagnostic in issue 20225 (closed 2026-03-08 as client prefix mutation), per hybrid-and-sliding-window-attention-kv-cache-rewinding.md. [source]
- The mismatch print sits in the same block as the `pos_min_thold` restore logic that PR 24110 touched, so it is maintained with the checkpoint code. [source]
- No log when `n_past == 0`: a client that changes the very first tokens (a timestamp at the top of the system prompt) shows no mismatch dump, only a cold prefill. [source]
- The dump prints on every request with a partial prefix, including benign appends where `n_past` equals the old length; for those, the `old` side ends at the marker and the `new` side continues. [source]
- The dump needs the warn level only; the surrounding `n_past = ..., pos_min = ..., n_swa = ...` line and `Checking checkpoint with [a, b] against thold...` lines are trace-level (`SLT_TRC`) and need `-lv 4` per the existing dossiers. [source]
- A tight N_DIFF window can hide the real cause when a client change sits a few tokens before `n_past`; widen it before concluding the template is stable. [source]
- None: sources agree the flag shows what changed, not why. The "why" side, client mutation versus server bug, is argued in 24055 and 20225 and is already in other dossiers. [source]
- Whether a future build adds a CLI flag or `/props` field; none seen. [source]
- Whether launchd or llama-swap pass the variable to a child process by default; the existing dossiers cover launchd env only for Ollama. [source]
- The slots-debug dump runs only when `n_past > 0`. [source]
- With N_DIFF unset the dump covers two token positions, `n_past` and `n_past + 1`. [source]
- The dump prints before the checkpoint restore decision. [source]
- Setting `LLAMA_SERVER_SLOTS_DEBUG` adds the full prompt text and generated text to each slot object returned by the slots endpoint. [source]
- Both variables are parsed with `atoi`, so a non-numeric value reads as 0 and disables the feature silently. [source]
- A 23013 log shows the line sequence `n_past = 85490, slot.prompt.tokens.size() = 89584, seq_id = 0, pos_min = 89583, n_swa = 0` followed by `Checking checkpoint with [89323, 89323] against 85490...` repeated per checkpoint. [source]
- That 23013 log means every checkpoint had `pos_min` above the 85,490 threshold, so none qualified and the whole prompt was reprocessed. [source]
- The slot selection line reads `selected slot by LCP similarity, sim_best = ..., f_keep = ...`, printed before the dump. [source]
Children
- No children recorded.