<!-- llms-explorer concept facts · https://llms-explorer.com/tree/llama-server-slots-debug-mismatch-logging-workfl/ · pack 2026-10-05 · ~1541 tokens -->

# LLAMA_SERVER_SLOTS_DEBUG mismatch logging workflow

> At load the server logs `LLAMA_SERVER_SLOTS_DEBUG = <n>` as a warning when it is non-zero, and `LLAMA_SERVER_SLOTS_N_DIFF = <n>` when that one is non-zero; seeing those two lines is the check that the environment reached the process.

Parent: [Mac local LLMs: Prompt cache and persistent KV](https://llms-explorer.com/tree/mac-local-llms-prompt-cache-and-persistent-kv/) · 1 facets · 25 facts · page: https://llms-explorer.com/tree/llama-server-slots-debug-mismatch-logging-workfl/

## Facts

- At load the server logs `LLAMA_SERVER_SLOTS_DEBUG = <n>` as a warning when it is non-zero, and `LLAMA_SERVER_SLOTS_N_DIFF = <n>` when that one is non-zero; seeing those two lines is the check that the environment reached the process. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-context.cpp)
- The dump runs only when `n_past > 0 && n_past <= slot.prompt.n_tokens()`. If the new prompt shares no prefix with the slot (`n_past == 0`), nothing is printed. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-context.cpp)
- Window: `np0 = max(n_past - N_DIFF, 0)` and `np1 = min(n_past + N_DIFF + 2, min(old size, new size))`. With N_DIFF unset (0) the dump shows only token indices `n_past` and `n_past + 1`. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-context.cpp)
- Each token prints as its text piece (`common_token_to_piece`) and as an 8-wide id; a media placeholder prints as `[mtmd]`; a ` | ` marker is inserted at index `n_past`, the first differing position. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-context.cpp)
- Four warnings follow in order: `old: ... <text>`, `new: ... <text>`, then the old token ids and the new token ids. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-context.cpp)
- The dump comes before the checkpoint search, so it prints even when a checkpoint then restores successfully. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-context.cpp)
- With the debug flag on, the slot list endpoint serializes with `only_metrics = false`, so each slot adds `prompt` (the detokenized task tokens) and `generated` (the generated text, or the copy saved at slot release). — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-context.cpp)
- At release the server copies `slot.generated_text` into `debug_generated_text` only when the flag is set. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-context.cpp)
- ggerganov named the flag as the diagnostic in issue 20225 (closed 2026-03-08 as client prefix mutation), per hybrid-and-sliding-window-attention-kv-cache-rewinding.md. — source: `asserted`
- The mismatch print sits in the same block as the `pos_min_thold` restore logic that PR 24110 touched, so it is maintained with the checkpoint code. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-context.cpp)
- No log when `n_past == 0`: a client that changes the very first tokens (a timestamp at the top of the system prompt) shows no mismatch dump, only a cold prefill. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-context.cpp)
- The dump prints on every request with a partial prefix, including benign appends where `n_past` equals the old length; for those, the `old` side ends at the marker and the `new` side continues. — source: `asserted`
- The dump needs the warn level only; the surrounding `n_past = ..., pos_min = ..., n_swa = ...` line and `Checking checkpoint with [a, b] against thold...` lines are trace-level (`SLT_TRC`) and need `-lv 4` per the existing dossiers. — source: `asserted`
- A tight N_DIFF window can hide the real cause when a client change sits a few tokens before `n_past`; widen it before concluding the template is stable. — source: `asserted`
- None: sources agree the flag shows what changed, not why. The "why" side, client mutation versus server bug, is argued in 24055 and 20225 and is already in other dossiers. — source: `asserted`
- Whether a future build adds a CLI flag or `/props` field; none seen. — source: `asserted`
- Whether launchd or llama-swap pass the variable to a child process by default; the existing dossiers cover launchd env only for Ollama. — source: `asserted`
- The slots-debug dump runs only when `n_past > 0`. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-context.cpp)
- With N_DIFF unset the dump covers two token positions, `n_past` and `n_past + 1`. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-context.cpp)
- The dump prints before the checkpoint restore decision. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-context.cpp)
- Setting `LLAMA_SERVER_SLOTS_DEBUG` adds the full prompt text and generated text to each slot object returned by the slots endpoint. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-context.cpp)
- Both variables are parsed with `atoi`, so a non-numeric value reads as 0 and disables the feature silently. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-context.cpp)
- A 23013 log shows the line sequence `n_past = 85490, slot.prompt.tokens.size() = 89584, seq_id = 0, pos_min = 89583, n_swa = 0` followed by `Checking checkpoint with [89323, 89323] against 85490...` repeated per checkpoint. — [source](https://github.com/ggml-org/llama.cpp/issues/23013)
- That 23013 log means every checkpoint had `pos_min` above the 85,490 threshold, so none qualified and the whole prompt was reprocessed. — [source](https://github.com/ggml-org/llama.cpp/issues/23013)
- The slot selection line reads `selected slot by LCP similarity, sim_best = ..., f_keep = ...`, printed before the dump. — [source](https://github.com/ggml-org/llama.cpp/issues/24714)
