llama.cpp message_spans chat-template user-boundary detection for checkpoints
Parent: Mac local LLMs: llama.cpp internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Template analysis builds the delimiters: in the differential autoparser branch of `common_chat_templates_apply`, `assistant_start` and `user_start` strings found by the autoparser become role delimiters (`delimiters.add(COMMON_CHAT_ROLE_ASSISTANT, ...)`, `delimiters.add(COMMON_CHAT_ROLE_USER, ......
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Template analysis builds the delimiters: in the differential autoparser branch of `common_chat_templates_apply`, `assistant_start` and `user_start` strings found by the autoparser become role delimiters (`delimiters.add(COMMON_CHAT_ROLE_ASSISTANT, ...)`, `delimiters.add(COMMON_CHAT_ROLE_USER, ...)`), stored in `common_chat_params::message_delimiters`. [source]
- The OAI-compat conversion writes them into the request as `llama_params["message_delimiters"]`, a JSON array of `{role, delimiter}` objects. [source]
- The completion handler reads `message_delimiters` back (`common_chat_msg_delimiters_parse`), tokenizes each delimiter string with `common_tokenize(vocab, delimiter, false, true)`, and calls `task.tokens.find_message_spans(delimiters)` per task. [source]
- `split()` scans the prompt tokens left to right, tests every delimiter's token sequence at each index with `std::equal`, records the first match as (role, index), and closes each span at the next match; media chunks are jumped over through a skip map built from `map_idx_to_media`. [source]
- The result is `common_chat_msg_spans` with `is_user_start(pos)` (exact position match on a user span) and `last_user_message_pos()` (last user span, or -1). [source]
- `message_spans` lives in the task params struct (`server-task.h`, field `message_spans`). [source]
- In `update_slots`, the batch-fill loop breaks right after pushing a token when `do_checkpoint && spans.is_user_start(n_tokens)` and either the position is the last user message, or no checkpoint exists, or the position is more than `checkpoint_min_step` past the newest checkpoint. [source]
- A checkpoint is then created if the batch started at a user start or is near prompt end (`task.n_tokens() < n_tokens + n_ubatch`); other mid-prompt checkpoints are skipped. [source]
- PR 22929 first computed the position by text (byte offsets); PR 24176 moved to a direct token scan, and PR 25420 (merged 2026-07-09) made the batch split respect min-step. [source]
- Delimiters travel in the request JSON, so a client or proxy that builds a raw `/completion` call can supply its own `message_delimiters`. [source]
- Only the differential-autoparser path sets delimiters in the code read; the other branches that return early (a plain generation-prompt template, specialized templates) leave `message_delimiters` empty, so no user boundary is found and only near-end checkpoints are made. [source]
- A delimiter that tokenizes differently in context than alone (BPE merge across a newline) would not match `std::equal` on the token ids. [source]
- Multimodal prompts: spans skip media chunks, and `do_checkpoint` is turned off for the batch that follows an mtmd chunk. [source]
- The two end-of-prompt checkpoints (offsets `4 + n_ubatch` and `4`) exist independently of spans; see llama-cpp-slot-state-trimming-when-the-template.md. [source]
- None found between sources on the data path; the earlier PR text says "extract from chat templates" and the code agrees, with the template analysis, not the Jinja source, supplying the strings. [source]
- Which shipped model templates yield non-empty `user_start` (Qwen3.5/3.6 are reported to benefit; no list exists). [source]
- Whether a tool-result message (role tool) ever starts a span that is treated as a user boundary: only `COMMON_CHAT_ROLE_USER` counts in `is_user_start`. [source]
- `common_chat_msg_delimiters` holds a list of (role, delimiter string, tokens) entries and exposes `tokenize(vocab)`, `split(tokens, skips)` and `to_json()`. [source]
- `common_chat_params` has a `message_delimiters` member. [source]
- The role enum includes system, assistant, user, tool and unknown, and `find_message_spans` is a `server_tokens` method. [source]
- `server_tokens::find_message_spans` builds a skip map from media chunk start indices to `mtmd_input_chunk_get_n_tokens` and delegates to `delims.split`. [source]
- The completion path attaches `message_spans` to every task right after `eval_llama_cmpl_schema`. [source]
- The last span is closed with an unknown-role sentinel at `tokens.size()`, so the final message length runs to the end of the prompt. [source]
- Checkpoint creation requires `slot.task->type == SERVER_TASK_TYPE_COMPLETION` and a model whose seq_rm type is FULL or RS, or `n_swa > 0`. [source]
- A user-start break does not by itself create a checkpoint; the checkpoint also needs `pos_min >= 0`, no mtmd chunk in the batch, and (empty list, last user message, or beyond min-step). [source]
- PR 25420 states the old logic "breaks on every user message start, which does not respect --checkpoint-min-step". [source]
Children
- No children recorded.