Chat-template and tool-declaration determinism as a cache-hit prerequisite
Parent: Mac local LLMs: Prompt cache and persistent KV · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Ollama stores tool properties and tool-call arguments in insertion-ordered maps (`ToolPropertiesMap`, `ToolCallFunctionArguments`). A renderer that loops with `.All()` therefore emits keys in the order the client's JSON had them.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Ollama stores tool properties and tool-call arguments in insertion-ordered maps (`ToolPropertiesMap`, `ToolCallFunctionArguments`). A renderer that loops with `.All()` therefore emits keys in the order the client's JSON had them. [source]
- The Qwen3.5 renderer writes the tool block first, inside the leading system turn, and serialises each tool with `marshalWithSpaces`. [source]
- `marshalWithSpaces` runs `json.Marshal` and then inserts a space after every `:` and `,` that lies outside a string, to mimic Python `json.dumps` spacing. [source]
- The Qwen3.5 renderer renders an assistant `<think>` block only for assistant turns after the last real user query, unless the Qwen3.8 variant flag `alwaysRenderAssistantThinkBlock` is set. [source]
- The Qwen3-Coder renderer takes the first system message and drops later ones, and adds the default system text "You are Qwen, a helpful AI assistant that can interact with a computer to solve tasks." when tools exist but no system message does. [source]
- The Qwen3-Coder renderer emits extra schema keys (anything except `type` and `description` on a property, and anything except `type` and `properties` on the parameters object) through `renderAdditionalKeys`. That helper marshals the object, unmarshals it into a `map[string]any`, and ranges over the map. [source]
- Ollama moved newer models from Go `TEMPLATE` strings to registered renderer and parser pairs; the qwen3-coder renderer carries a TODO to match Python `json.dumps` spacing, and the later Qwen3.5 renderer added `marshalWithSpaces` to do so. [source]
- llama.cpp issue 26974 (2026-08-12) showed template render time growing quadratically with tool count; PR 27034 merged 2026-08-14 with byte-identical prompts. [source]
- Go does not define an iteration order for `range` over a map. In `renderAdditionalKeys`, a property that carries two or more extra keys (for example `enum` and `default`) can therefore render those keys in a different order on different requests of the same logical tool set, which breaks the cached prefix. This is inferred from the code and the Go language rule; no issue reporting it was found. [source]
- The Qwen3.5 renderer has no such map loop for tool schemas, so the hazard is specific to the qwen3-coder renderer path. [source]
- Because the tool block precedes the system text and all messages, adding, removing or reordering one tool, or an MCP server reconnecting with its tools in a new order, changes bytes at the start of the prompt and forces a full re-prefill. [source]
- Insertion-ordered storage makes determinism depend on the client: a client that re-serialises `parameters.properties` with sorted or hash-ordered keys changes the rendered bytes even though the schema is logically equal. [source]
- Template render time is paid on every request, including full cache hits. At 20 tools of 20 properties the bundled Command-R tool-use template took 0.43 s on llama.cpp before the fix, and 9.2 s at 50 tools of 50 properties. [source]
- None found. The two engines choose different layers for determinism: Ollama fixes key order in Go structures, llama.cpp relies on the Jinja template and `common_json` insertion order (see llama-cpp-jinja-capability-probes-chat-template.md). [source]
- Whether `renderAdditionalKeys` ordering has produced an observable cache miss on a real qwen3-coder deployment; no measurement exists. [source]
- Whether Ollama's prefix reuse compares tokens or text, which decides how much one reordered key costs. [source]
- Ollama `ToolPropertiesMap` and `ToolCallFunctionArguments` keep insertion order through an ordered map and expose it through `All()`. [source]
- Ollama `marshalWithSpaces` inserts a space after `:` and `,` outside string values after `json.Marshal`. [source]
- The Ollama Qwen3.5 renderer renders each tool with `marshalWithSpaces` inside the first system turn, before the system content. [source]
- The Ollama Qwen3.5 renderer renders assistant think blocks only for turns after the last user query unless the Qwen3.8 variant is used. [source]
- The Ollama Qwen3-Coder renderer ranges over a Go `map[string]any` when it emits extra tool-schema keys. [source]
- The Ollama Qwen3-Coder renderer keeps only the first system message and injects a default system message when tools exist without one. [source]
- llama.cpp issue 26974 measured chat-template render time of 0.43 s for 20 tools by 20 properties and 9.3 s for 50 by 50 on the Command-R tool-use template. [source]
- llama.cpp PR 27034 (merged 2026-08-14) fixed two quadratic terms in `gather_string_parts` and kept prompts byte-identical. [source]
- In Ollama's qwen3-coder renderer, two requests with the same tools can render extra schema keys in different orders. [source]
- A change in tool order or tool set alters the first system turn and invalidates the whole prefix cache on Qwen renderers. [source]
Children
- No children recorded.