llama_n_ctx_seq semantics with and without unified KV
Parent: Mac local LLMs: llama.cpp internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
llama-server-props-n-ctx-as-the-per-slot-window.md lists as an open question whether `llama_n_ctx_seq` divides the pool by slot count when KV is not unified; this file answers it from `src/llama-context.cpp`.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- llama-server-props-n-ctx-as-the-per-slot-window.md lists as an open question whether `llama_n_ctx_seq` divides the pool by slot count when KV is not unified; this file answers it from `src/llama-context.cpp`. [source]
- With `kv_unified` true, context setup sets `cparams.n_ctx_seq = cparams.n_ctx`, so every sequence's window is the whole pool (after n_ctx is padded to 256). [source]
- With `kv_unified` false, it sets `n_ctx_seq = n_ctx / n_seq_max` and pads it with `GGML_PAD(..., 256)`, so the pool is split evenly across sequences. [source]
- Without unified KV, `n_ctx_seq == 0` after division throws `n_ctx_seq == 0`. [source]
- Without unified KV, if `n_ctx != n_ctx_seq * n_seq_max` the context rewrites `n_ctx` to that product and logs `n_ctx is not divisible by n_seq_max - rounding down to %u`. [source]
- The context logs `n_ctx_seq (%u) < n_ctx_train (%u) -- the full capacity of the model will not be utilized` and warns `n_ctx_seq (%u) > n_ctx_train (%u) -- possible training context overflow` for the opposite case. [source]
- `llama_n_ctx_seq(ctx)` is declared in `include/llama.h` and returns `ctx->n_ctx_seq()`, the stored `cparams.n_ctx_seq`. [source]
- `llama_context_params.kv_unified` is documented as using a unified buffer across the input sequences when computing attention and defaults to false in `llama_context_default_params`. [source]
- Output buffer sizing uses `LLAMA_MAX_SEQ` instead of `n_seq_max` when `kv_unified` is true, a separate effect of the same flag. [source]
Children
- No children recorded.