XGrammar structural tags for Gemma 4 and Harmony tool calls
Parent: Mac local LLMs: Chat templates, reasoning and tool calling · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Issue 587 (opened 2026-04-09) gives the Gemma 4 thinking format as `<|channel>thought\n...<channel|>` with `<|channel>` token ID 100 and `<channel|>` token ID 101, and says `enable_thinking=False` makes the template pre-close an empty channel `<|channel>thought\n<channel|>`.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Issue 587 (opened 2026-04-09) gives the Gemma 4 thinking format as `<|channel>thought\n...<channel|>` with `<|channel>` token ID 100 and `<channel|>` token ID 101, and says `enable_thinking=False` makes the template pre-close an empty channel `<|channel>thought\n<channel|>`. [source]
- Issue 587 argues that without a structural tag, grammar-constrained decoding and thinking are mutually exclusive for Gemma 4: constraints from the start break thinking tokens, and disabling thinking loses reasoning. [source]
- PR 588 (merged 2026-04-13, commit 2f4a0c5) registered `gemma4` and describes the format as optional thinking, optional content, `<|tool_call>call:name{json}<tool_call|>` calls and optional `<turn|>`. [source]
- PR 609 removed `force_empty_reasoning`, renamed model keys (`qwen` to `qwen_3`, `qwen_coder` to `qwen_3_coder`, `glm47` to `glm_4_7`, `gemma4` to `gemma_4`) and added a `qwen_3_5` key for Qwen3.5 and 3.6 XML tool calls with an empty reasoning part. [source]
- Under PR 609 `reasoning=False` forces an empty reasoning block for kimi and qwen_3_5, omits the reasoning section for llama, qwen_3_coder, deepseek_r1, deepseek_v3_2, deepseek_v4, glm_4_7, gemma_4 and minimax, strips the `<think>` prefix for qwen_3, and omits the analysis channel for harmony. [source]
- PR 609's review notes that renaming model keys is a breaking change and suggests aliases for backward compatibility. [source]
- PR 923 says the `gemma_4` builtin "has been unregistered with a 'we will support it later' TODO since #609". [source]
- PR 923 adds a `"gemma"` JSON-schema style that emits unquoted object keys and strings delimited by the single `<|"|>` token with no escape sequences at every nesting level, and re-enables `gemma_4` with auto, required and forced tool choice, the `<|channel>thought` reasoning prefix and `parallel_tool_calls`. [source]
- PR 923 continues PR 684 (by iwanhae), which was closed unmerged after the converter was rewritten to build grammar trees directly (PR 745), and says engines fall back to unconstrained decoding for Gemma 4 tool calls because earlier converters cannot express `<|"|>`-delimited strings. [source]
- PR 923 registers `<|"|>` as an exclusion so string patterns and length-bounded strings cannot contain it, makes `const` or `enum` literals that spell it unsatisfiable, and excludes declared property names from `additionalProperties` keys. [source]
- PR 923's tests: 59 converter cases, 612 builtin structural-tag tests, and 34 alignment tests over vendored e2b and 31b Gemma 4 templates with thinking on and off; its known limits are `:`-containing key patterns and picojson re-serialisation of numbers inside `const` objects. [source]
- Issue 587's timeline shows PR 923's commit referencing it dated 2026-09-24, about one week before the fetch. [source]
- vLLM issue 39130 (2026-04-06, v0.19.0, closed) reports `--reasoning-parser gemma4` silently disabling structured output through xgrammar when `enable_thinking=false`, because the reasoning end token is never detected. [source]
- vLLM issue 53363 (2026-08-22, open) reports `tool_choice: "required"` not enforced with the gemma4 tool parser: a prose-only reply came back with `finish_reason: "tool_calls"` and an empty `tool_calls` array. [source]
- Inferred: any Mac runtime that builds on XGrammar can constrain Gemma 4 tool calls only after taking a release that includes PR 923; before it the `gemma_4` key was absent. [source]
- Inferred: llama.cpp's Gemma 4 handling does not use XGrammar and so is unaffected by this change. [source]
Children
- No children recorded.