llama.cpp reasoning-preserve vs Qwen preserve_thinking template key aliasing
Parent: Mac local LLMs: Chat templates, reasoning and tool calling · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
2026-06-28: PR #25105 "jinja, chat: add --reasoning-preserve flag" merged (supersedes #25079), commit b3fed31b.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- 2026-06-28: PR #25105 "jinja, chat: add --reasoning-preserve flag" merged (supersedes #25079), commit b3fed31b. [source]
- 2026-07-22: #25999 fixes the reasoning-preserve variable for DS4. [source]
- 2026-09-16: PR #28987 (StepFun templates honour preserve_thinking) opened and closed unmerged; shows the alias only helps templates that read one of the four names. [source]
- Default on: any template that passes the probe gets preserved history unless `--no-reasoning-preserve`. [source]
- A template that reads none of the four names (or fails the probe) is unaffected; the flag is a silent no-op. [source]
- Qwen3.8 template (line 117) treats `preserve_thinking is undefined` as true, so it preserves even without the flag. [source]
- Qwen3.6 template needs `is true`; without the alias it would drop history thinking. With a current build the alias supplies it. [source]
- Builds older than b3fed31b (before 2026-06-28) lack the alias: use `--chat-template-kwargs '{"preserve_thinking":true}'`. [source]
- `--chat-template-kwargs` containing `preserve_reasoning` logs a deprecation warning but still works. [source]
- The #54 complaint is likely from a pre-b3fed31b build, a non-llama.cpp frontend, or templates not matching the probe; unverified. [source]
- No runtime test of stock Qwen3.6 GGUF was run here; conclusion rests on source reading. [source]
- Whether GGUF-embedded templates differ from the HF chat_template.jinja (Unsloth edits). [source]
- Per-request `chat_template_kwargs` is accepted in the request body and merged over server-wide `--chat-template-args` (JSON string), then passed to `tokenizer.apply_chat_template`. Added in PR #829 (2026-02-03). No `--reasoning-preserve` equivalent exists; set `{"preserve_thinking": true}` yourself. [source]
- llama.cpp (any build >= 2026-06-28): `llama-server --jinja -m <qwen3.6.gguf>` (preserve is on by default). Explicit: add `--reasoning-preserve`. Verify: `curl localhost:8080/props` shows `chat_template_caps.supports_preserve_reasoning: true`. [source]
- Belt and braces / older builds: add `--chat-template-kwargs '{"preserve_thinking":true}'`. [source]
- mlx_lm.server: `mlx_lm.server --model <m> --chat-template-args '{"preserve_thinking":true}'`, or per request `"chat_template_kwargs":{"preserve_thinking":true}`. [source]
- arg.cpp sets default_template_kwargs["preserve_reasoning"]="true" when the user did not specify it [source]
- `--reasoning-preserve` / `--no-reasoning-preserve` write "true"/"false" to preserve_reasoning, env LLAMA_ARG_REASONING_PRESERVE, for server, completion and cli [source]
- chat.cpp calls jinja::caps_apply_preserve_reasoning when the input has a boolean preserve_reasoning [source]
- caps_apply_preserve_reasoning sets preserve_thinking=enabled and clear_thinking, truncate_history_thinking, drop_thinking=!enabled [source]
- supports_preserve_reasoning is detected by rendering a history whose non-latest assistant turn carries a reasoning placeholder and checking it appears in output [source]
- global_from_json runs after caps_apply_preserve_reasoning, so an explicit preserve_thinking in kwargs overrides the alias [source]
- Qwen3.6-27B chat_template.jinja line 101 gates historical think blocks on `preserve_thinking is defined and preserve_thinking is true` [source]
- Qwen3.8-27B template line 117 preserves when `preserve_thinking is undefined`, so default is preserve [source]
- PR #25105 "jinja, chat: add --reasoning-preserve flag" merged 2026-06-28 [source]
- caps.cpp history lists the flag commit b3fed31b on 2026-06-28 and a DS4 variable fix on 2026-07-22 [source]
- PR #28987 (StepFun templates honour preserve_thinking) was closed unmerged [source]
- mlx_lm.server merges request chat_template_kwargs over server chat_template_args and passes them to apply_chat_template [source]
- mlx_lm.server defines `--chat-template-args` as a JSON string for apply_chat_template [source]
- mlx-lm added server chat_template_kwargs support in PR #829 on 2026-02-03 [source]
- Stock Qwen3.6 GGUFs get preserved thinking from the generic flag on builds from 2026-06-28 onward (inferred from source, not run) [source]
- mlx_lm.server has no generic preserve flag; users must pass preserve_thinking explicitly [source]
Corrections and disagreements
- Existing dossier and froggeric #54 commenters: the two keys are unaliased, stock Qwen ignores the generic flag. Source code: llama.cpp aliases them internally (caps_apply_preserve_reasoning). CONTRADICTS: qwen3-6-preserve-thinking-chat-template-flag-and-prompt-cache-reuse.md (claim that "a template that does not alias the two will not see the generic flag" and the open question "no source tested"). [source]
Children
- No children recorded.