<!-- llms-explorer concept facts · https://llms-explorer.com/tree/ollama-qwen3-8-reasoning-effort-instructions-inj/ · pack 2026-10-05 · ~1896 tokens -->

# Ollama Qwen3.8 reasoning-effort instructions injected into the system turn

> Ollama's `qwen3.8` renderer is `newQwen38Renderer()`, a variant of `Qwen35Renderer` with `isThinking` true, `alwaysRenderAssistantThinkBlock` and `emitEmptyThinkOnNoThink`, registered in `model/renderers/renderer.go` under `"qwen3.8"`.

Parent: [Mac local LLMs: Chat templates, reasoning and tool calling](https://llms-explorer.com/tree/mac-local-llms-chat-templates-reasoning-and-tool-calling/) · 1 facets · 25 facts · page: https://llms-explorer.com/tree/ollama-qwen3-8-reasoning-effort-instructions-inj/

## Facts

- Ollama's `qwen3.8` renderer is `newQwen38Renderer()`, a variant of `Qwen35Renderer` with `isThinking` true, `alwaysRenderAssistantThinkBlock` and `emitEmptyThinkOnNoThink`, registered in `model/renderers/renderer.go` under `"qwen3.8"`. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/renderers/qwen35.go)
- A stock `qwen3.8:27b` logs `selected=renderer_parser renderer=qwen3.8 parser=qwen3.5`, so Qwen3.8 reuses the Qwen3.5 parser. — [source](https://api.github.com/repos/ollama/ollama/issues/18632)
- The xhigh sentence is "Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer." — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/renderers/qwen35.go)
- The low sentence is "Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration." — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/renderers/qwen35.go)
- `qwen38ReasoningInstructions` returns the xhigh sentence when `think` is nil, the low sentence for `"low"`, nothing for `"medium"` or `false`, the xhigh sentence for `"xhigh"` and for any other valid string, and the error "invalid thinking value %v" for an invalid value. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/renderers/qwen35.go)
- A renderer test named "boolean true uses API medium effort" renders `think: true` like Jinja `reasoning_effort="medium"`, so `think: true` adds no sentence while an unset `think` adds the xhigh sentence. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/renderers/qwen38_test.go)
- With tools, the system turn starts with the sentence, then a blank line, then "# Tools" and the function list, then the user's system text after another blank line. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/renderers/qwen35.go)
- Without tools, a system turn holds the sentence, a blank line and the system text; with a sentence and no system message the system turn is the sentence alone; with no sentence and empty system text the qwen3.8 renderer writes no system turn. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/renderers/qwen35.go)
- `normalizeQwen38Messages` merges every `system` and `developer` message into one leading system message joined by blank lines, then drops them from the conversation order; a system or developer message with images is rejected. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/renderers/qwen35.go)
- `validateMessages` fails with "no messages provided", "system message cannot contain images" or "no user query found in messages", the last when every user message is a `<tool_response>` block. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/renderers/qwen35.go)
- An Ollama user reported "qwen 3.8 reports error during query: ... no user query found in messages (status code: 500" in a streaming chat. — [source](https://api.github.com/search/issues?q=%22exceed_context_size_error%22+claude&per_page=20)
- On main the renderer's `Thinking()` metadata is values `[false, "low", "medium", "xhigh"]` with default `"xhigh"`. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/renderers/qwen35.go)
- Issue 18632 shows `/api/show` for `qwen3.8:27b` (Ollama 0.34.4) listing `[false, "low", "medium", "xhigh"]` with default `medium`, and `/api/chat` answering HTTP 200 with the same bytes for `"medium"`, `"high"`, `"max"` and `"banana"`. — [source](https://api.github.com/repos/ollama/ollama/issues/18632)
- In that test `"xhigh"` produced 45,367 thinking characters and 12,185 eval tokens against 13,252 characters and 4,585 tokens for `"medium"` (temperature 0, seed 1), so xhigh costs about 2.7 times the tokens. — [source](https://api.github.com/repos/ollama/ollama/issues/18632)
- The 18632 reporter notes Ollama's validator documents the `think` vocabulary as `high`, `medium`, `low`, `max`, `true` or `false`, which excludes `xhigh`, and that issue 18091 was answered with "Use `max`". — [source](https://api.github.com/repos/ollama/ollama/issues/18632)
- Issue 17906 reproduced `ollama launch claude --model 'hf.co/orcarouter/Qwen3.8-27B-Uncensored-GGUF:Q4_K_M' -- --effort xhigh` failing with HTTP 500 on `POST /v1/messages?beta=true` and "Jinja Exception: Unexpected reasoning effort high. Supported types are xhigh (default), medium, and low." (Claude Code 2.1.177 and 2.1.238, Ollama 0.32.13 and 0.32.15). — [source](https://api.github.com/repos/ollama/ollama/issues/17906)
- `--effort medium` avoided that failure because `medium` passed through unchanged. — [source](https://api.github.com/repos/ollama/ollama/issues/17906)
- PR 17917, which stopped converting `xhigh` to `high` in `FromMessagesRequest`, was closed on 2026-08-27 without being merged. — [source](https://api.github.com/repos/ollama/ollama/pulls/17917)
- Issue 18766 (Ollama 0.35.0) reports that on a Qwen3.8 GGUF whose embedded template has `resolved_reasoning_effort`, `/api/chat` accepts `think: "low"`, `"medium"` and `"high"` but behaves like `true`, running at the template default xhigh. — [source](https://api.github.com/repos/ollama/ollama/issues/18766)
- In agent use `think: "low"` produced 1,809 thinking tokens for a one-line chat message, while rendering the template by hand with `reasoning_effort="low"` and sending it with `raw: true` gave 10 to 66 thinking tokens. — [source](https://api.github.com/repos/ollama/ollama/issues/18766)
- PR 17745 (merged 2026-08-14) states Qwen3.8 keeps the Qwen3.5 architecture and parser but its template adds reasoning-effort and preserved-thinking semantics, and detects those markers during safetensors import to select the `qwen3.8` renderer. — [source](https://api.github.com/repos/ollama/ollama/pulls/17745)
- PR 18786 (open, 2026-10-04) selects `qwen3.8` and `qwen3.5` for `qwen35` and `qwen35moe` GGUF imports when the embedded template contains both `resolved_reasoning_effort` and `preserve_thinking`, and says it does not change effort-name validation (issue 18632). — [source](https://api.github.com/repos/ollama/ollama/issues/18786)
- The Ollama library page for qwen3.8 says thinking is on by default, depth is tuned with `reasoning_effort`, and `preserve_thinking` keeps reasoning from earlier messages. — [source](https://ollama.com/library/qwen3.8)
- Inferred: on a native `ollama pull qwen3.8` model, Claude Code effort `low` and `medium` map to the low sentence and no sentence, and `xhigh` maps to the xhigh sentence; `high` and `max` give the xhigh sentence on main but `medium` behaviour on 0.34.4. — source: `asserted`
- Inferred: a client that omits `think` and one that sends `think: true` get different system turns (sentence versus none), so their prompt prefixes do not share a KV cache. — source: `asserted`
