Mac local LLMs: Chat templates, reasoning and tool calling
Parent: Running LLM models locally on a Mac · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
mlx_lm.server infers the tool parser from marker strings in the chat template (`_infer_tool_parser()`); no known marker means raw text stays in `content`. `tool_parser_type` in tokenizer_config.json is an override, not the only route.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Parser and template wiring (who parses what)
- mlx_lm.server infers the tool parser from marker strings in the chat template (`_infer_tool_parser()`); no known marker means raw text stays in `content`. `tool_parser_type` in tokenizer_config.json is an override, not the only route. [source]
- Ollama newer models use Go RENDERER/PARSER named in the manifest, not the GGUF Jinja. llama.cpp (PR #18675, ~2026-03-05) builds a PEG parser from the template. LM Studio parser is closed-source. vLLM: `--reasoning-parser` plus `--tool-call-parser`; `--tool-call-parser openai` for gpt-oss. [source]
- Qwen3.5/3.6 use the Qwen3-Coder XML format; every runtime shipped a fragile path in Feb 2026 (Ollama #14493 mis-wired to Hermes JSON, fixed by a dedicated Qwen35Parser). Qwen3.8 reuses parser `qwen3.5`. [source]
Tool call lands inside thinking
- Symptom: `tool_calls: []`, XML in `reasoning_content`, `finish_reason: stop` (LM Studio, llama.cpp #20837, vLLM #39056 on 0.18-0.19). Workaround: `enable_thinking=false` via `--chat-template-kwargs '{"enable_thinking":false}'`. [source]
- Fixed on current main: vLLM single-engine parser recovers calls in reasoning (PR 39055 closed unmerged 2026-06-23; engine in 0.24); Ollama Qwen35Parser injects `</think>` before `<tool_call>`; mlx_lm.server state machine collects them. Rapid-MLX still discards them. [source]
- llama-server has `-rea on|off|auto` and per-request `reasoning_effort` none; "none" sets enable_thinking=false, other strings are forwarded verbatim. [source]
- llama.cpp reasoning-budget sampler arms for any template with think tags; default budget INT_MAX (log: `reasoning-budget: activated, budget=2147483647 tokens`) caused hangs and KV spill. forge uses `--reasoning-budget 0` (Gemma 4 31B: 340 s vs 100 s); that also removes lazy-grammar protection inside thinking. [source]
llama.cpp auto-parser (template probing, chat-diff-analyzer.cpp)
- `analyze_template` probes reasoning, then content, then tools by rendering the template with sentinel messages and diffing. If the template throws on a probe, it logs and silently yields no reasoning, plain content or no tool format; check the debug dump. [source]
- Thinking detection compares `enable_thinking` false vs true; if both sides differ it falls back to tail anchoring (last 64 chars). Modes: reasoning NONE/TAG_BASED/TOOLS_ONLY; tool format JSON_NATIVE/TAG_WITH_JSON/TAG_WITH_TAGGED. [source]
- 12 hardcoded workaround lambdas patch known templates (Granite 3.3, Cohere, Functionary 3.1, R1-Distill-Qwen, Nemotron Nano v2, Fireworks v2, Solar Open, Apriel 1.6, Laguna, Bailing V3); an unlisted odd template relies on the diff path. [source]
Preserve thinking and cache reuse
- Default Qwen templates drop past reasoning, so the cached prefix diverges; hybrid delta-net models (Qwen3.5/3.6) cannot trim recurrent state, so a mismatch costs far more. [source]
- llama.cpp: `--chat-template-kwargs '{"preserve_thinking":true}'` (needs `--jinja`; env `LLAMA_ARG_CHAT_TEMPLATE_KWARGS`). Builds from b3fed31b (2026-06-28) have `--reasoning-preserve`, default on, which aliases to `preserve_reasoning`; templates reading none of the names ignore it silently. [source]
- The client must echo `reasoning_content` back, or the flag has nothing to render. Per request: `extra_body={"chat_template_kwargs":{"preserve_thinking":True}}`. [source]
- mlx_lm.server: `--chat-template-args` JSON plus per-request `chat_template_kwargs`. oMLX turns preserve on when the template text contains `preserve_thinking`; `forced_ct_kwargs` silently drops request overrides; on render `TypeError` it retries without all user kwargs. [source]
- Qwen3.8: `reasoning_effort` default `xhigh` adds a system sentence at offset 0 (changing effort re-prefills everything); only xhigh/medium/low valid, else `Unexpected reasoning effort`; llama-server forwards "high"/"max" into that error. [source]
- Templates trim whitespace and re-serialise args, so re-render differs from generated tokens; measure with `/apply-template`, `/tokenize`, `return_tokens:true` and `cache_n`. [source]
- Gemma 4: strip prior thoughts between ordinary turns, never between calls inside one turn; expect a cache break. gpt-oss Harmony: drop `analysis` after a `final`, keep it across tool calls. [source]
Error strings and fixes
- `System message must be at the beginning` (Qwen3.8 templates, LM Studio 0.4.15+2, Ollama `system message must be at the beginning`): Claude Code sends role:system mid-`messages`. Ollama 0.32.14 passes it through (PR 17757); llama.cpp needs a patched template (discussion 27281); LiteLLM #40693 open. [source]
- Ollama `500 no user query found in messages`: first request ~36k tokens exceeds default num_ctx 32768; raise num_ctx. [source]
- mlx-lm `ValueError: No function provided.` (qwen3_coder.py line 110, issue #905) on empty `<tool_call>`; `Received 333 parameters not in model` needs PR #928. [source]
- llama.cpp 500 `Callee is not a function: got Undefined (hint: 'items')` on no-argument calls (GLM-4.7, b7756). Codex 400 `'type' of tool must be 'function'`: fixed by llama.cpp #23041 skipping non-function tools. [source]
Decision rules
- Prevention order: native template tool calling, grammar constraint, fewer tools, repair last; a malformed rate above ~1 in 20 means fix config. Parse-only repair can raise risk. [source]
- vLLM tool grammar: envelope always constrained; arguments only if `strict: true` or `--tool-strict-level parameter`. [source]
- Reasoning field: vLLM removed `reasoning_content` (PR #33402, 2026-01-30) for `reasoning`; llama-server emits `reasoning_content`. [source]
- Ollama `think` string must match `/api/show` `thinking.values`; structured outputs unsupported on Ollama Cloud. [source]
Corrections and open questions
- Missing `count_tokens` is optional in Claude Code (estimate, no error); earlier "404 stalls" claim is doubtful; Ollama tool_use survival is one reporter (Rallo, 2026-04-04), unresolved. [source]
- Ollama `qwen tool call parsing failed ... element <parameter> closed by </function>` (0.32.1) drops the turn with HTTP 200. Open: does LiteLLM `modify_params` help local backends; is LiteLLM PR 40726 merged. [source]
Corrections and disagreements
- CONTRADICTS local-llm-server-as-a-coding-agent-backend-on-mac.md (line 16, "tool_parser_type is needed"): mlx-lm infers the parser from the template, so the tokenizer_config.json key is an explicit override rather than the only mechanism [source]
- CONTRADICTS qwen3-6-preserve-thinking-chat-template-flag-and-prompt-cache-reuse.md (line "mlx_lm.server: no source found"): mlx_lm.server has a `--chat-template-args` JSON flag and a per-request `chat_template_kwargs` field merged over it [source]
- CONTRADICTS tool-call-parser-and-chat-template-mismatch.md (line 79, "hard enable_thinking switch is not exposed by llama.cpp CLI flags"): current llama-server has `-rea on|off|auto` and per-request `reasoning_effort` none [source]
- CONTRADICTS local-llm-server-as-a-coding-agent-backend-on-mac.md (Ollama >= 0.14 serves /v1/messages natively): the Rallo article (2026-04-04) says Ollama tool_use blocks do not survive translation; it is one reporter and a single date, so treat it as an unresolved conflict. [source]
- Existing dossier and froggeric #54 commenters: the two keys are unaliased, stock Qwen ignores the generic flag. Source code: llama.cpp aliases them internally (caps_apply_preserve_reasoning). CONTRADICTS: qwen3-6-preserve-thinking-chat-template-flag-and-prompt-cache-reuse.md (claim that "a template that does not alias the two will not see the generic flag" and the open question "no source tested"). [source]
- This CONTRADICTS the causal framing in ollama-anthropic-adapter-tool-use-block-fidelity.md ("count_tokens 404 ... stalls"): Claude Code's protocol page lists `/v1/messages/count_tokens` as optional and says a gateway lacking it gets a character-based estimate, so `/context` shows approximate counts, with no error. [source]
- CONTRADICTS: reasoning-parser-versus-tool-parser-interaction.md line 45, which lists Ollama among runtimes where a tool call inside thinking lands in the reasoning field with `tool_calls` empty. On current main `Qwen35Parser` detects `<tool_call>` inside thinking, injects `</think>` before it, and parses the call. The symptom applies to builds before that code. [source]
- CONTRADICTS: reasoning-parser-versus-tool-parser-interaction.md, which states for vLLM that "reasoning extraction runs first and the tool parser reads only content; a call left in reasoning is lost" and that PR 39055's merge state was not confirmed. PR 39055 closed unmerged on 2026-06-23, and on current main a call inside reasoning is recovered by grammar. The statement holds only for vLLM before the engine parser (0.18 to 0.23). [source]
- CONTRADICTS: tool-call-parser-and-chat-template-mismatch.md line 74-75, which treats `qwen3_coder` and `qwen3_xml` as different parsers with different robustness. On main they share one engine. Third-party comparisons of their regex versus expat internals describe the older code. [source]
Concepts in this cluster
- Tool-call parser and chat-template mismatch [source]
- Qwen3.6 preserve_thinking chat-template flag and prompt-cache reuse [source]
- Reasoning-parser versus tool-parser interaction [source]
- Tool-call repair proxies and self-healing parsers [source]
- Tool-calling benchmarks for local models [source]
- Anthropic-compatible local endpoints and mid-conversation role system messages [source]
- Gemma 4 thought stripping in agentic turns [source]
- Ollama structured outputs on thinking models [source]
- Qwen chat template whitespace normalisation of tool-call and think blocks [source]
- Reasoning field naming across servers [source]
- gpt-oss Harmony reasoning replay across tool calls [source]
- llama.cpp reasoning-budget sampler and lazy grammar trigger gating [source]
- llama.cpp reasoning-preserve vs Qwen preserve_thinking template key aliasing [source]
- oMLX per-model chat template kwargs and cache-key construction [source]
- LiteLLM Anthropic-to-Chat-Completions translation bugs for Claude Code [source]
- Qwen 3.8 chat-template fixes and reasoning-level mapping [source]
- Codex Responses API compatibility on local servers [source]
- Ollama Anthropic adapter tool_use block fidelity [source]
- OpenAI gpt-oss compatibility test script for verifying server implementations [source]
- Reasoning-block trailing newline parser bug causing agent loops [source]
- Token-level generation-versus-render drift measurement harness [source]
- XGrammar structural tags for Gemma 4 and Harmony tool calls [source]
- Ollama count_tokens 404 stall with Claude Code [source]
- Ollama native tool-call parser extraction for Qwen function XML [source]
- Parse-then-render byte round-trip tests for reasoning turns [source]
- llama.cpp Responses SSE event ordering and reasoning items [source]
- vLLM XML tool_call emitted inside think block loses tool calls [source]
- Ollama Qwen3.8 reasoning-effort instructions injected into the system turn [source]
- vLLM parser engine grammar (vllm/parser) and structural-tag strict tool calling [source]
- llama.cpp autoparser workaround lambdas in chat-diff-analyzer.cpp [source]
- llama.cpp autoparser analyze_reasoning, analyze_content and analyze_tools diff p [source]
Children
- Ollama Anthropic adapter tool_use block fidelity
- Ollama count_tokens 404 stall with Claude Code
- Ollama native tool-call parser extraction for Qwen function XML
- Ollama Qwen3.8 reasoning-effort instructions injected into the system turn
- Ollama structured outputs on thinking models
- oMLX per-model chat template kwargs and cache-key construction
- OpenAI gpt-oss compatibility test script for verifying server implementations
- Parse-then-render byte round-trip tests for reasoning turns
- Qwen 3.8 chat-template fixes and reasoning-level mapping
- Qwen chat template whitespace normalisation of tool-call and think blocks
- Qwen3.6 preserve_thinking chat-template flag and prompt-cache reuse
- Reasoning-block trailing newline parser bug causing agent loops
- Reasoning field naming across servers
- Reasoning-parser versus tool-parser interaction
- Token-level generation-versus-render drift measurement harness
- Tool-call parser and chat-template mismatch
- Tool-call repair proxies and self-healing parsers
- Tool-calling benchmarks for local models
- vLLM parser engine grammar (vllm/parser) and structural-tag strict tool calling
- vLLM XML tool_call emitted inside think block loses tool calls
- XGrammar structural tags for Gemma 4 and Harmony tool calls
- Anthropic-compatible local endpoints and mid-conversation role system messages
- Codex Responses API compatibility on local servers
- Gemma 4 thought stripping in agentic turns
- gpt-oss Harmony reasoning replay across tool calls
- LiteLLM Anthropic-to-Chat-Completions translation bugs for Claude Code
- llama.cpp autoparser analyze_reasoning, analyze_content and analyze_tools diff p
- llama.cpp autoparser workaround lambdas in chat-diff-analyzer.cpp
- llama.cpp reasoning-budget sampler and lazy grammar trigger gating
- llama.cpp reasoning-preserve vs Qwen preserve_thinking template key aliasing
- llama.cpp Responses SSE event ordering and reasoning items