<!-- llms-explorer concept facts · https://llms-explorer.com/tree/llama-cpp-reasoning-preserve-vs-qwen-preserve-th/ · pack 2026-10-05 · ~1621 tokens -->

# llama.cpp reasoning-preserve vs Qwen preserve_thinking template key aliasing

> 2026-06-28: PR #25105 "jinja, chat: add --reasoning-preserve flag" merged (supersedes #25079), commit b3fed31b.

Parent: [Mac local LLMs: Chat templates, reasoning and tool calling](https://llms-explorer.com/tree/mac-local-llms-chat-templates-reasoning-and-tool-calling/) · 2 facets · 33 facts · page: https://llms-explorer.com/tree/llama-cpp-reasoning-preserve-vs-qwen-preserve-th/

## Facts

- 2026-06-28: PR #25105 "jinja, chat: add --reasoning-preserve flag" merged (supersedes #25079), commit b3fed31b. — source: `asserted`
- 2026-07-22: #25999 fixes the reasoning-preserve variable for DS4. — source: `asserted`
- 2026-09-16: PR #28987 (StepFun templates honour preserve_thinking) opened and closed unmerged; shows the alias only helps templates that read one of the four names. — source: `asserted`
- Default on: any template that passes the probe gets preserved history unless `--no-reasoning-preserve`. — source: `asserted`
- A template that reads none of the four names (or fails the probe) is unaffected; the flag is a silent no-op. — source: `asserted`
- Qwen3.8 template (line 117) treats `preserve_thinking is undefined` as true, so it preserves even without the flag. — source: `asserted`
- Qwen3.6 template needs `is true`; without the alias it would drop history thinking. With a current build the alias supplies it. — source: `asserted`
- Builds older than b3fed31b (before 2026-06-28) lack the alias: use `--chat-template-kwargs '{"preserve_thinking":true}'`. — source: `asserted`
- `--chat-template-kwargs` containing `preserve_reasoning` logs a deprecation warning but still works. — source: `asserted`
- The #54 complaint is likely from a pre-b3fed31b build, a non-llama.cpp frontend, or templates not matching the probe; unverified. — source: `asserted`
- No runtime test of stock Qwen3.6 GGUF was run here; conclusion rests on source reading. — source: `asserted`
- Whether GGUF-embedded templates differ from the HF chat_template.jinja (Unsloth edits). — source: `asserted`
- Per-request `chat_template_kwargs` is accepted in the request body and merged over server-wide `--chat-template-args` (JSON string), then passed to `tokenizer.apply_chat_template`. Added in PR #829 (2026-02-03). No `--reasoning-preserve` equivalent exists; set `{"preserve_thinking": true}` yourself. — source: `asserted`
- llama.cpp (any build >= 2026-06-28): `llama-server --jinja -m <qwen3.6.gguf>` (preserve is on by default). Explicit: add `--reasoning-preserve`. Verify: `curl localhost:8080/props` shows `chat_template_caps.supports_preserve_reasoning: true`. — source: `asserted`
- Belt and braces / older builds: add `--chat-template-kwargs '{"preserve_thinking":true}'`. — source: `asserted`
- mlx_lm.server: `mlx_lm.server --model <m> --chat-template-args '{"preserve_thinking":true}'`, or per request `"chat_template_kwargs":{"preserve_thinking":true}`. — source: `asserted`
- arg.cpp sets default_template_kwargs["preserve_reasoning"]="true" when the user did not specify it — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/common/arg.cpp)
- `--reasoning-preserve` / `--no-reasoning-preserve` write "true"/"false" to preserve_reasoning, env LLAMA_ARG_REASONING_PRESERVE, for server, completion and cli — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/common/arg.cpp)
- chat.cpp calls jinja::caps_apply_preserve_reasoning when the input has a boolean preserve_reasoning — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/common/chat.cpp)
- caps_apply_preserve_reasoning sets preserve_thinking=enabled and clear_thinking, truncate_history_thinking, drop_thinking=!enabled — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/common/jinja/caps.cpp)
- supports_preserve_reasoning is detected by rendering a history whose non-latest assistant turn carries a reasoning placeholder and checking it appears in output — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/common/jinja/caps.cpp)
- global_from_json runs after caps_apply_preserve_reasoning, so an explicit preserve_thinking in kwargs overrides the alias — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/common/chat.cpp)
- Qwen3.6-27B chat_template.jinja line 101 gates historical think blocks on `preserve_thinking is defined and preserve_thinking is true` — [source](https://huggingface.co/Qwen/Qwen3.6-27B/raw/main/chat_template.jinja)
- Qwen3.8-27B template line 117 preserves when `preserve_thinking is undefined`, so default is preserve — [source](https://huggingface.co/Qwen/Qwen3.8-27B/raw/main/chat_template.jinja)
- PR #25105 "jinja, chat: add --reasoning-preserve flag" merged 2026-06-28 — [source](https://api.github.com/repos/ggml-org/llama.cpp/pulls/25105)
- caps.cpp history lists the flag commit b3fed31b on 2026-06-28 and a DS4 variable fix on 2026-07-22 — [source](https://api.github.com/repos/ggml-org/llama.cpp/commits?path=common/jinja/caps.cpp&per_page=30)
- PR #28987 (StepFun templates honour preserve_thinking) was closed unmerged — [source](https://api.github.com/repos/ggml-org/llama.cpp/pulls/28987)
- mlx_lm.server merges request chat_template_kwargs over server chat_template_args and passes them to apply_chat_template — [source](https://raw.githubusercontent.com/ml-explore/mlx-lm/main/mlx_lm/server.py)
- mlx_lm.server defines `--chat-template-args` as a JSON string for apply_chat_template — [source](https://raw.githubusercontent.com/ml-explore/mlx-lm/main/mlx_lm/server.py)
- mlx-lm added server chat_template_kwargs support in PR #829 on 2026-02-03 — [source](https://api.github.com/repos/ml-explore/mlx-lm/commits?path=mlx_lm/server.py&per_page=60)
- Stock Qwen3.6 GGUFs get preserved thinking from the generic flag on builds from 2026-06-28 onward (inferred from source, not run) — source: `asserted`
- mlx_lm.server has no generic preserve flag; users must pass preserve_thinking explicitly — source: `asserted`

## Corrections and disagreements

- Existing dossier and froggeric #54 commenters: the two keys are unaliased, stock Qwen ignores the generic flag. Source code: llama.cpp aliases them internally (caps_apply_preserve_reasoning). CONTRADICTS: qwen3-6-preserve-thinking-chat-template-flag-and-prompt-cache-reuse.md (claim that "a template that does not alias the two will not see the generic flag" and the open question "no source tested"). — source: `asserted`
