llama-server /props n_ctx as the per-slot window to advertise to clients
Parent: Mac local LLMs: llama.cpp internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
`n_ctx_slot()` in `tools/server/server-context.cpp` returns `llama_n_ctx_seq(ctx_tgt)`, capped to `kv_unified_per_slot` when that is above zero, and then capped to `llama_model_n_ctx_train(model_tgt)`.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- `n_ctx_slot()` in `tools/server/server-context.cpp` returns `llama_n_ctx_seq(ctx_tgt)`, capped to `kv_unified_per_slot` when that is above zero, and then capped to `llama_model_n_ctx_train(model_tgt)`. [source]
- The `/props` response builds `default_generation_settings` with `n_ctx` set to `meta.slot_n_ctx`, which is `impl->n_ctx_slot()`, so the advertised number already includes both caps. [source]
- Each slot's `n_ctx` is assigned from `n_ctx_slot()` at slot initialisation, and the startup log line reads `initializing, n_slots = %d, n_ctx_slot = %d, kv_unified = '%s'`. [source]
- When `--kv-unified-per-slot` is smaller than the per-slot pool capacity, the server logs `capping per-slot context (%d) to --kv-unified-per-slot (%d)`. [source]
- When `--kv-unified-per-slot` exceeds the per-slot pool capacity, the server logs a warning that the cap has no effect and says to raise the pool with `-c` or unset `-c` so the pool is sized to `n_parallel * kv_unified_per_slot`. [source]
- When the slot context exceeds the model's training context the server logs `the slot context (%d) exceeds the training context of the model (%d) - capping`. [source]
- `/props` also returns `total_slots` (the `--parallel` value), `model_path`, `chat_template`, `chat_template_caps`, `modalities`, `media_marker`, `build_info` and `is_sleeping`; `POST /props` works only when the server runs with `--props`. [source]
- In router mode the README says GET endpoints such as `/props` take the target in a `model` query parameter, for example `/props?model=ggml-org%2Fgemma-3-4b-it-GGUF%3AQ4_K_M`. [source]
- `/props` also reports a sleeping state, and the README ties it to the sleeping-on-idle feature. [source]
- A model-metadata JSON built in `server-context.cpp` carries `n_ctx_train` (from `meta.model_n_ctx_train`) beside the slot `n_ctx`, so the training length and the usable slot window are separate numbers. [source]
Children
- No children recorded.