llama.cpp Jinja input marking is_input special-token injection defence
Parent: Mac local LLMs: llama.cpp internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
`common/chat-auto-parser.h` defaults `mark_input = true` and `common/chat.cpp` passes it to `global_from_json`, so the Jinja side marks input strings.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- `common/chat-auto-parser.h` defaults `mark_input = true` and `common/chat.cpp` passes it to `global_from_json`, so the Jinja side marks input strings. [source]
- `tools/server/` has no reference to `is_input` or `mark_input`, and all `tokenize_input_prompts` calls pass `parse_special = true`, so a user message containing a ChatML header tokenizes to the same ids as a real system turn. [source]
- On current master the multimodal path still builds `mtmd_input_text` with `parse_special` set to true. [source]
- The wiring plan is issue 28249 (opened 2026-09-02, open, assigned to the maintainer). Its checklist is: add `mtmd_tokenize_from_parts` (done), refactor `chat.cpp` to output parts instead of one string, wire it to the server. [source]
- PR 28250 added `mtmd_tokenize_from_parts()` for per-segment tokenization, where one segment may use `parse_special` and another may not; it merged on 2026-09-02. [source]
- 2026-06-09, issue 24382: `{"role":"user","content":"A<|turn>B"}` on a Gemma 4 GGUF produced token 105 (`<|turn>`) from user text. Maintainer ngxson said the Jinja infrastructure came after `chat.cpp`, and that regex replacements `chat.cpp` applies to some templates make sanitisation tricky. [source]
- 2026-07-29, issue 26273: a request for `--escape-special-in-input` described models that copy injected special tokens into reasoning and confuse the auto-parser, producing generation that ends mid-thought, reasoning leaking into the answer, or malformed tool calls. It closed as stale on 2026-09-12. [source]
- 2026-08-03, issue 26532: the missing server wiring was reported; maintainer CISC said it is still on ngxson's list but "not trivial (if at all 100% feasible) so not high priority". [source]
- 2026-09-02: issue 24382 was closed as replaced by issue 28249. [source]
- 2026-09-21: a commenter on issue 26273 reproduced on Qwen3.8-27B (Unsloth Q4_K_S, `--jinja`): a tool result containing the literal `<|im_end|>` stopped generation at that point and the text after it was never reproduced; building the string from fragments `"<" + "|im_end|" + ">"` was fine. [source]
- Agent tool output routinely contains tokenizer metadata, configs and code with special-token strings, so injection through tool results is more likely than through user typing. [source]
- Input marking cannot protect a token the model itself writes. Maintainer aldehir (issue 27661, 2026-08-24): Qwen 3 models do not mark the `</think>` token as special, so even sanitised input tokenizes to the same id, and a model that writes `</think>` inside its reasoning ends the reasoning early in the parser; the only fix would break streaming. [source]
- Turning escaping on would change the token ids of any prompt that contains a special-token string, so cached prefixes holding such text would stop matching on the first request after the change. [source]
- Template-built special tokens such as `'<|' + role + '|>'` stop working once marking is consumed (README caveat held in the existing dossier). [source]
- Python `transformers` has no equivalent: `apply_chat_template(tokenize=True)` lets content forge turns, and a split-special flag on the two-step path also strips the template's own delimiters. This is the issue 26532 author's claim. [source]
- Feasibility. The 26532 author and the 26273 requester treat a flag as small work (the latter attached a Codex-written patch adding `--escape-special-in-input`). CISC and ngxson call full coverage hard because `chat.cpp` rewrites template text with regexes before rendering. Both positions are stated in the threads and not resolved. [source]
- When the `chat.cpp` parts refactor in issue 28249 lands, and whether input-marked tokenization will be default-on. [source]
- Whether the llama-server completion endpoint will honour parts for non-multimodal prompts, since the first merged piece is in the multimodal library. [source]
- On the 26532 report (commit 3581ba0), `mark_input` defaults to true in `common/chat-auto-parser.h` and `common/chat.cpp` passes it to `global_from_json`. [source]
- On the 26532 report, `tools/server/` has zero references to `is_input` or `mark_input`. [source]
- On the 26532 report, all four `tokenize_input_prompts` calls in `server-context.cpp` pass `parse_special = true`. [source]
- On master read 2026-10-04, `server-common.cpp` builds the multimodal `mtmd_input_text` with `parse_special` true. [source]
- Issue 28249 "wiring up jinja input marking" is open and lists two unchecked items: refactor `chat.cpp` to output parts, and wire it to the server. [source]
- PR 28250 `mtmd_tokenize_from_parts()` merged 2026-09-02 as the first step of input marking. [source]
- Issue 24382 showed user text `A<|turn>B` tokenizing to Gemma 4 special token 105. [source]
- ngxson said regex replacements in `chat.cpp` for certain templates make sanitisation of input tricky. [source]
- CISC called the server wiring "not trivial (if at all 100% feasible)" and low priority on 2026-08-04. [source]
- A Qwen3.8-27B tool result containing literal `<|im_end|>` ended generation at that position under llama-server `--jinja`. [source]
- aldehir stated Qwen 3 models do not mark `</think>` as special, so input sanitisation could not disambiguate a model-written `</think>` from the real one. [source]
- Existing open question answered: input marking is produced by the Jinja engine by default but not enabled in llama-server tokenization as of 2026-10-04. [source]
- Enabling input-escape tokenization would invalidate cached prefixes that contain special-token strings. [source]
Children
- No children recorded.