<!-- llms-explorer concept facts · https://llms-explorer.com/tree/llama-n-ctx-seq-semantics-with-and-without-unifi/ · pack 2026-10-05 · ~674 tokens -->

# llama_n_ctx_seq semantics with and without unified KV

> llama-server-props-n-ctx-as-the-per-slot-window.md lists as an open question whether `llama_n_ctx_seq` divides the pool by slot count when KV is not unified; this file answers it from `src/llama-context.cpp`.

Parent: [Mac local LLMs: llama.cpp internals](https://llms-explorer.com/tree/mac-local-llms-llama-cpp-internals/) · 1 facets · 9 facts · page: https://llms-explorer.com/tree/llama-n-ctx-seq-semantics-with-and-without-unifi/

## Facts

- llama-server-props-n-ctx-as-the-per-slot-window.md lists as an open question whether `llama_n_ctx_seq` divides the pool by slot count when KV is not unified; this file answers it from `src/llama-context.cpp`. — source: `asserted`
- With `kv_unified` true, context setup sets `cparams.n_ctx_seq = cparams.n_ctx`, so every sequence's window is the whole pool (after n_ctx is padded to 256). — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/src/llama-context.cpp)
- With `kv_unified` false, it sets `n_ctx_seq = n_ctx / n_seq_max` and pads it with `GGML_PAD(..., 256)`, so the pool is split evenly across sequences. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/src/llama-context.cpp)
- Without unified KV, `n_ctx_seq == 0` after division throws `n_ctx_seq == 0`. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/src/llama-context.cpp)
- Without unified KV, if `n_ctx != n_ctx_seq * n_seq_max` the context rewrites `n_ctx` to that product and logs `n_ctx is not divisible by n_seq_max - rounding down to %u`. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/src/llama-context.cpp)
- The context logs `n_ctx_seq (%u) < n_ctx_train (%u) -- the full capacity of the model will not be utilized` and warns `n_ctx_seq (%u) > n_ctx_train (%u) -- possible training context overflow` for the opposite case. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/src/llama-context.cpp)
- `llama_n_ctx_seq(ctx)` is declared in `include/llama.h` and returns `ctx->n_ctx_seq()`, the stored `cparams.n_ctx_seq`. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/include/llama.h)
- `llama_context_params.kv_unified` is documented as using a unified buffer across the input sequences when computing attention and defaults to false in `llama_context_default_params`. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/src/llama-context.cpp)
- Output buffer sizing uses `LLAMA_MAX_SEQ` instead of `n_seq_max` when `kv_unified` is true, a separate effect of the same flag. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/src/llama-context.cpp)
