llama.cpp reasoning-preserve flag and template support detection
Parent: Mac local LLMs: llama.cpp internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Ollama has no preserve flag. Its Qwen3.5 renderer decides per turn: an assistant turn gets a `<think>` block only when thinking is on and the turn comes after the last real user query, or when the renderer variant forces it.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Ollama has no preserve flag. Its Qwen3.5 renderer decides per turn: an assistant turn gets a `<think>` block only when thinking is on and the turn comes after the last real user query, or when the renderer variant forces it. [source]
- The Qwen3.8 variant sets `alwaysRenderAssistantThinkBlock`, so it renders a think block for every assistant turn and behaves like llama.cpp preserve-on. [source]
- No new dated events found beyond those in the existing dossier. [source]
- In the Ollama Qwen3.5 renderer, a user query that follows a finished tool loop moves the last-query index forward, so the earlier think blocks vanish from the prompt and the cached prefix breaks at the first of them. llama.cpp with preserve on avoids this for templates that read one of the four alias names. [source]
- The `reasoning_content` that Ollama feeds back is passed through `strings.TrimSpace`, so whitespace the model generated around its thinking is not reproduced. [source]
- None added. The existing open questions (runtime test of a stock Qwen3.6 GGUF, GGUF versus Hugging Face template differences) remain open. [source]
- This concept is fully covered by llama-cpp-reasoning-preserve-vs-qwen-preserve-thinking-aliasing.md and llama-cpp-jinja-capability-probes-chat-template.md; no llama.cpp claim absent from them was found. [source]
- The Ollama Qwen3.5 renderer renders assistant think blocks only for turns after the last user query unless the Qwen3.8 variant flag is set. [source]
- The Ollama Qwen3.8 renderer variant always renders an assistant think block. [source]
- Ollama trims whitespace from message thinking and content when it re-renders an assistant turn. [source]
Corrections and disagreements
- None found. The existing CONTRADICTS note in llama-cpp-reasoning-preserve-vs-qwen-preserve-thinking-aliasing.md stands; nothing read here changes it. [source]
Children
- No children recorded.