<!-- llms-explorer concept facts · https://llms-explorer.com/tree/llama-server-props-n-ctx-as-the-per-slot-window/ · pack 2026-10-05 · ~839 tokens -->

# llama-server /props n_ctx as the per-slot window to advertise to clients

> `n_ctx_slot()` in `tools/server/server-context.cpp` returns `llama_n_ctx_seq(ctx_tgt)`, capped to `kv_unified_per_slot` when that is above zero, and then capped to `llama_model_n_ctx_train(model_tgt)`.

Parent: [Mac local LLMs: llama.cpp internals](https://llms-explorer.com/tree/mac-local-llms-llama-cpp-internals/) · 1 facets · 10 facts · page: https://llms-explorer.com/tree/llama-server-props-n-ctx-as-the-per-slot-window/

## Facts

- `n_ctx_slot()` in `tools/server/server-context.cpp` returns `llama_n_ctx_seq(ctx_tgt)`, capped to `kv_unified_per_slot` when that is above zero, and then capped to `llama_model_n_ctx_train(model_tgt)`. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-context.cpp)
- The `/props` response builds `default_generation_settings` with `n_ctx` set to `meta.slot_n_ctx`, which is `impl->n_ctx_slot()`, so the advertised number already includes both caps. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-context.cpp)
- Each slot's `n_ctx` is assigned from `n_ctx_slot()` at slot initialisation, and the startup log line reads `initializing, n_slots = %d, n_ctx_slot = %d, kv_unified = '%s'`. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-context.cpp)
- When `--kv-unified-per-slot` is smaller than the per-slot pool capacity, the server logs `capping per-slot context (%d) to --kv-unified-per-slot (%d)`. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-context.cpp)
- When `--kv-unified-per-slot` exceeds the per-slot pool capacity, the server logs a warning that the cap has no effect and says to raise the pool with `-c` or unset `-c` so the pool is sized to `n_parallel * kv_unified_per_slot`. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-context.cpp)
- When the slot context exceeds the model's training context the server logs `the slot context (%d) exceeds the training context of the model (%d) - capping`. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-context.cpp)
- `/props` also returns `total_slots` (the `--parallel` value), `model_path`, `chat_template`, `chat_template_caps`, `modalities`, `media_marker`, `build_info` and `is_sleeping`; `POST /props` works only when the server runs with `--props`. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/README.md)
- In router mode the README says GET endpoints such as `/props` take the target in a `model` query parameter, for example `/props?model=ggml-org%2Fgemma-3-4b-it-GGUF%3AQ4_K_M`. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/README.md)
- `/props` also reports a sleeping state, and the README ties it to the sleeping-on-idle feature. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/README.md)
- A model-metadata JSON built in `server-context.cpp` carries `n_ctx_train` (from `meta.model_n_ctx_train`) beside the slot `n_ctx`, so the training length and the usable slot window are separate numbers. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-context.cpp)
