Ollama GGUF import renderer detection (qwen3.8 renderer selection)
Parent: Mac local LLMs: Ollama internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
This ANSWERS the open point in ollama-qwen3-8-reasoning-effort-instructions-inj.md about how PR 18786 selects the renderer: it adds a `case "qwen35", "qwen35moe":` to the GGUF auto-detect switch in `server/create.go` and reads `layer.GGUF.String("tokenizer.chat_template")`.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- This ANSWERS the open point in ollama-qwen3-8-reasoning-effort-instructions-inj.md about how PR 18786 selects the renderer: it adds a `case "qwen35", "qwen35moe":` to the GGUF auto-detect switch in `server/create.go` and reads `layer.GGUF.String("tokenizer.chat_template")`. [source]
- PR 18786 sets renderer `qwen3.8` and parser `qwen3.5` only when the embedded chat template contains both `resolved_reasoning_effort` and `preserve_thinking`. [source]
- The PR states those two markers are the same ones Ollama already uses to choose the renderer for safetensors imports. [source]
- `createModel` fills `Renderer` and `Parser` with `cmp.Or(config.X, ...)`, so a value set in the request or Modelfile is kept and only the empty one is detected. [source]
- The detect block runs only when `config.Renderer == "" || config.Parser == ""` and only for layers with media type `application/vnd.ollama.image.model`. [source]
- On main before the PR, the GGUF architecture switch in `server/create.go` covers `gemma4` (renderer legacy alias, parser `gemma4`, stop `<turn|>`), `laguna` (renderer and parser `laguna`) and `nemotron_h`, `nemotron_h_moe`, `nemotron_h_omni` (renderer and parser `nemotron-3-nano`), with no Qwen case. [source]
- The `gemma4` case also sets the `stop` parameter to `<turn|>` when the request has none. [source]
- The PR adds nine model-creation regression cases: dense and MoE, explicit pair and partial overrides, absent template, each missing marker, and an unrelated architecture. [source]
- The PR author reports that reverting the production change makes four detection and partial-override cases fail while the controls pass, and that no live model inference or full server suite was run. [source]
- The renderer registry in `model/renderers/renderer.go` names `qwen3-coder`, `qwen3-vl-instruct`, `qwen3-vl-thinking`, `qwen3.5` and `qwen3.8`; the parser registry in `model/parsers/parsers.go` names `qwen3`, `qwen3-thinking`, `qwen3.5`, `qwen3-coder`, `qwen3-vl-instruct` and `qwen3-vl-thinking`, and has no `qwen3.8` entry. [source]
- The renderer registry returns the error `unknown renderer %q` for a name that is not registered. [source]
Children
- No children recorded.