<!-- llms-explorer concept facts · https://llms-explorer.com/tree/llama-cpp-reasoning-preserve-flag-and-template-s/ · pack 2026-10-05 · ~738 tokens -->

# llama.cpp reasoning-preserve flag and template support detection

> Ollama has no preserve flag. Its Qwen3.5 renderer decides per turn: an assistant turn gets a `<think>` block only when thinking is on and the turn comes after the last real user query, or when the renderer variant forces it.

Parent: [Mac local LLMs: llama.cpp internals](https://llms-explorer.com/tree/mac-local-llms-llama-cpp-internals/) · 2 facets · 11 facts · page: https://llms-explorer.com/tree/llama-cpp-reasoning-preserve-flag-and-template-s/

## Facts

- Ollama has no preserve flag. Its Qwen3.5 renderer decides per turn: an assistant turn gets a `<think>` block only when thinking is on and the turn comes after the last real user query, or when the renderer variant forces it. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/renderers/qwen35.go)
- The Qwen3.8 variant sets `alwaysRenderAssistantThinkBlock`, so it renders a think block for every assistant turn and behaves like llama.cpp preserve-on. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/renderers/qwen35.go)
- No new dated events found beyond those in the existing dossier. — source: `asserted`
- In the Ollama Qwen3.5 renderer, a user query that follows a finished tool loop moves the last-query index forward, so the earlier think blocks vanish from the prompt and the cached prefix breaks at the first of them. llama.cpp with preserve on avoids this for templates that read one of the four alias names. — source: `asserted`
- The `reasoning_content` that Ollama feeds back is passed through `strings.TrimSpace`, so whitespace the model generated around its thinking is not reproduced. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/renderers/qwen35.go)
- None added. The existing open questions (runtime test of a stock Qwen3.6 GGUF, GGUF versus Hugging Face template differences) remain open. — source: `asserted`
- This concept is fully covered by llama-cpp-reasoning-preserve-vs-qwen-preserve-thinking-aliasing.md and llama-cpp-jinja-capability-probes-chat-template.md; no llama.cpp claim absent from them was found. — source: `asserted`
- The Ollama Qwen3.5 renderer renders assistant think blocks only for turns after the last user query unless the Qwen3.8 variant flag is set. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/renderers/qwen35.go)
- The Ollama Qwen3.8 renderer variant always renders an assistant think block. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/renderers/qwen35.go)
- Ollama trims whitespace from message thinking and content when it re-renders an assistant turn. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/renderers/qwen35.go)

## Corrections and disagreements

- None found. The existing CONTRADICTS note in llama-cpp-reasoning-preserve-vs-qwen-preserve-thinking-aliasing.md stands; nothing read here changes it. — source: `asserted`
