llama.cpp jinja capability probes chat_template_caps
Parent: Mac local LLMs: llama.cpp internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Probes run with `ctx.is_get_stats = true`. Every value the template reads is tagged with `stats.used` and the list of operations applied to it (`test_is_string`, `selectattr`, `array_access`). Most probes infer a capability from "was this field read", not from the rendered text. Exceptions during...
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Probes run with `ctx.is_get_stats = true`. Every value the template reads is tagged with `stats.used` and the list of operations applied to it (`test_is_string`, `selectattr`, `array_access`). Most probes infer a capability from "was this field read", not from the rendered text. Exceptions during a probe are swallowed. [source]
- Probe order: typed content, system prompt, single tool call with object arguments, (string arguments, only if object arguments failed), parallel tool calls, preserve reasoning, reasoning effort. [source]
- Typed-content probe: one user message with the marker `STRING_MARKER`. Array-style access (`selectattr` or `array_access`) means `supports_typed_content`. A render failure, or array access whose output lacks the marker, clears `supports_string_content`. A template that tests `is string` triggers a second probe with an empty content array. [source]
- System probe: a system plus a user message. If the system content value was never read, `supports_system_role` is false. [source]
- Tool probe: user, assistant with one tool call (object arguments), tool result, assistant, user, with one tool definition. Unread `tools[0].function.name` clears `supports_tools`; unread `tool_calls` clears `supports_tool_calls`; a read of `arguments.arg` sets `supports_object_arguments`. If that fails, a string-arguments variant runs, and a render failure there clears both `supports_tool_calls` and `supports_tools`. [source]
- Parallel probe: two tool calls in one assistant message; unread `tool_calls[1].function` clears `supports_parallel_tool_calls`. [source]
- Reasoning-effort probe: sets `enable_thinking` true and `reasoning_effort` "low", then reads the stats of `reasoning_effort`. Unlike the preserve-reasoning probe, it does not look at the output, so any read of the variable counts. [source]
- `caps_apply_reasoning_effort` binds `reasoning_effort` and `reasoning_strength` to the same string value, so templates written for either name see the request value. [source]
- The preserve-reasoning probe is the only one judged on output text, because a template may read `reasoning_content` in an `if` test without printing it. [source]
- Workarounds keyed on the original (probed) caps, in `common_chat_templates_apply`: developer role mapped to system unless the source contains `<|channel|>`; system message merged into the next message with a single "\n" (or dropped, with a warning, when it is the only message) when `supports_system_role` is false; empty-string `content` added to tool-call messages when the template supports tool calls; JSON-string tool arguments parsed to objects when `supports_object_arguments`. [source]
- Content normaliser: a string-only template gets array content joined with "\n" between text parts and no newline around media markers; a typed-only template gets each string wrapped as `[{"type":"text","text":...}]`. [source]
- Per-request `chat_template_kwargs` are stored as JSON strings and parsed into `extra_context`; a non-empty string `reasoning_effort` in the merged input triggers `caps_apply_reasoning_effort` before `global_from_json`. [source]
- `common_json` (the wrapper llama.cpp now uses) keeps object keys in insertion order. A tool-argument string parsed by `func_args_not_string` therefore keeps the model's key order when a template iterates it. [source]
- The older minja-based `chat_template_caps` carried a narrower set; the nine-flag set with `supports_preserve_reasoning` and `supports_reasoning_effort` is on master as of 2026-10-04. Dates of each flag's introduction were not found in the sources read. [source]
- No flag covers acceptance of a non-leading system message. A template can pass the system probe (leading system is read) and still call `raise_exception('System message must be at the beginning.')` for a later one, as the Qwen 3.8 template does. [source]
- Qwen 3.8's template passes the reasoning-effort probe (effort "low" is valid) but rejects values such as "high" at render time; the probe cannot see the accepted vocabulary. [source]
- A template that supports tool calls but never prints tool definitions triggers the warning "Template supports tool calls but does not natively describe tools". [source]
- Because the typed/string normaliser rewrites content before rendering, a template that handles only arrays receives "\n"-joined text for multi-part string input; this differs from what a Python `apply_chat_template` call would render. Prefix tests done in Python do not see it. [source]
- Probes use synthetic messages with a user-only or short history, so history-dependent behaviour (empty think wrappers, whitespace) is not probed except for preserve reasoning. [source]
- Which release introduced `supports_reasoning_effort` and `supports_object_arguments` in the property output. [source]
- Whether `/props` `chat_template_caps` reflects the `--chat-template-file` override or the embedded template when both exist (chat.cpp returns the caps of `template_tool_use` if present, else `template_default`). [source]
- caps.cpp exposes nine booleans: supports_string_content, supports_typed_content, supports_tools, supports_tool_calls, supports_parallel_tool_calls, supports_system_role, supports_preserve_reasoning, supports_reasoning_effort, supports_object_arguments [source]
- Capability probes render the template with is_get_stats set and read per-value `stats.used` and applied-operation lists rather than only the output [source]
- Exceptions raised while a probe renders are caught and ignored [source]
- The typed-content probe uses the marker STRING_MARKER and treats selectattr or array_access on content as typed-content support [source]
- supports_system_role is false when a system message's content value is never read during the system probe [source]
- The tool probe runs a single call with object arguments first and a string-arguments variant only when object arguments are unsupported [source]
- supports_parallel_tool_calls is false when the second tool call in an assistant message is never read [source]
- supports_reasoning_effort is set from whether reasoning_effort (set to "low" with enable_thinking true) was read, without checking output [source]
- caps_apply_reasoning_effort sets both reasoning_effort and reasoning_strength to the same string [source]
- Only the preserve-reasoning probe is judged on rendered output, because reasoning_content may be read in an if test without being printed [source]
- chat.cpp merges a leading system message into the next message with "\n" when supports_system_role is false, and drops it with a warning when it is the only message [source]
- chat.cpp maps the developer role to system for every template whose source lacks "<|channel|>" [source]
- chat.cpp parses JSON-string tool arguments into objects when the template supports object arguments, and throws "Failed to parse tool call arguments as JSON" on invalid JSON [source]
- chat.cpp adds an empty-string content to tool-call messages that lack it when the template supports tool calls [source]
- chat.cpp converts content to string-only or typed-only form to match the probed caps, joining string-only text parts with "\n" [source]
- chat.cpp warns that a template supporting tool calls but not tool descriptions may produce bad results and suggests overriding the template [source]
- chat.cpp applies caps_apply_reasoning_effort when the merged input holds a non-empty string reasoning_effort [source]
- chat.cpp returns the caps map of the tool-use template when one exists, otherwise of the default template [source]
- common_json keeps object keys in the order they were added [source]
- No capability flag covers non-leading system messages, so a template that raises on them can still report supports_system_role true [source]
- The reasoning-effort probe cannot detect which effort names a template accepts [source]
Children
- No children recorded.