<!-- llms-explorer concept facts · https://llms-explorer.com/tree/ollama-gguf-import-renderer-detection-qwen3-8-re/ · pack 2026-10-05 · ~860 tokens -->

# Ollama GGUF import renderer detection (qwen3.8 renderer selection)

> This ANSWERS the open point in ollama-qwen3-8-reasoning-effort-instructions-inj.md about how PR 18786 selects the renderer: it adds a `case "qwen35", "qwen35moe":` to the GGUF auto-detect switch in `server/create.go` and reads `layer.GGUF.String("tokenizer.chat_template")`.

Parent: [Mac local LLMs: Ollama internals](https://llms-explorer.com/tree/mac-local-llms-ollama-internals/) · 1 facets · 11 facts · page: https://llms-explorer.com/tree/ollama-gguf-import-renderer-detection-qwen3-8-re/

## Facts

- This ANSWERS the open point in ollama-qwen3-8-reasoning-effort-instructions-inj.md about how PR 18786 selects the renderer: it adds a `case "qwen35", "qwen35moe":` to the GGUF auto-detect switch in `server/create.go` and reads `layer.GGUF.String("tokenizer.chat_template")`. — [source](https://github.com/ollama/ollama/pull/18786)
- PR 18786 sets renderer `qwen3.8` and parser `qwen3.5` only when the embedded chat template contains both `resolved_reasoning_effort` and `preserve_thinking`. — [source](https://github.com/ollama/ollama/pull/18786)
- The PR states those two markers are the same ones Ollama already uses to choose the renderer for safetensors imports. — [source](https://github.com/ollama/ollama/pull/18786)
- `createModel` fills `Renderer` and `Parser` with `cmp.Or(config.X, ...)`, so a value set in the request or Modelfile is kept and only the empty one is detected. — [source](https://raw.githubusercontent.com/ollama/ollama/main/server/create.go)
- The detect block runs only when `config.Renderer == "" || config.Parser == ""` and only for layers with media type `application/vnd.ollama.image.model`. — [source](https://raw.githubusercontent.com/ollama/ollama/main/server/create.go)
- On main before the PR, the GGUF architecture switch in `server/create.go` covers `gemma4` (renderer legacy alias, parser `gemma4`, stop `<turn|>`), `laguna` (renderer and parser `laguna`) and `nemotron_h`, `nemotron_h_moe`, `nemotron_h_omni` (renderer and parser `nemotron-3-nano`), with no Qwen case. — [source](https://raw.githubusercontent.com/ollama/ollama/main/server/create.go)
- The `gemma4` case also sets the `stop` parameter to `<turn|>` when the request has none. — [source](https://raw.githubusercontent.com/ollama/ollama/main/server/create.go)
- The PR adds nine model-creation regression cases: dense and MoE, explicit pair and partial overrides, absent template, each missing marker, and an unrelated architecture. — [source](https://github.com/ollama/ollama/pull/18786)
- The PR author reports that reverting the production change makes four detection and partial-override cases fail while the controls pass, and that no live model inference or full server suite was run. — [source](https://github.com/ollama/ollama/pull/18786)
- The renderer registry in `model/renderers/renderer.go` names `qwen3-coder`, `qwen3-vl-instruct`, `qwen3-vl-thinking`, `qwen3.5` and `qwen3.8`; the parser registry in `model/parsers/parsers.go` names `qwen3`, `qwen3-thinking`, `qwen3.5`, `qwen3-coder`, `qwen3-vl-instruct` and `qwen3-vl-thinking`, and has no `qwen3.8` entry. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/parsers/parsers.go)
- The renderer registry returns the error `unknown renderer %q` for a name that is not registered. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/renderers/renderer.go)
