Parse-then-render byte round-trip tests for reasoning turns
Parent: Mac local LLMs: Chat templates, reasoning and tool calling · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Ollama's Qwen3-Coder parser trims trailing whitespace before a `<tool_call>` tag and leading whitespace after `</tool_call>`, and strips one leading and one trailing newline from every parameter value.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Ollama's Qwen3-Coder parser trims trailing whitespace before a `<tool_call>` tag and leading whitespace after `</tool_call>`, and strips one leading and one trailing newline from every parameter value. [source]
- The Ollama renderer then writes a tool call as `\n<tool_call>\n<function=NAME>` plus, per argument, `\n<parameter=NAME>\n` VALUE `\n</parameter>`, so the removed newlines come back only if the model used exactly that layout. [source]
- Values are coerced by schema type on parse: `null` in any case becomes nil, boolean accepts `true` and `false` in any case, a number with no decimal part becomes an integer, and an invalid boolean becomes false. The renderer prints booleans and numbers with `%v`, strings raw, and maps and slices as compact JSON. [source]
- The Ollama Qwen3.5 renderer runs `strings.TrimSpace` on the stored thinking and on message content before writing `<think>\n` REASONING `\n</think>\n\n` CONTENT. [source]
- The Ollama Qwen3.5 parser eats whitespace after `</think>` and trims leading whitespace in the thinking, so generated whitespace in both places is dropped on parse. [source]
- llama.cpp's PEG parser tests stream the input one character at a time through `parser.parse(prefix, is_partial)`, with a UTF-8 safe truncation, and compare the accumulated message with the expectation. [source]
- `test_template_generation_prompt` in llama.cpp checks the generation prompt string a template adds for plain, content-continuation and reasoning-continuation cases. [source]
- Ollama's parser carries a comment that it follows a "reference implementation" for newline trimming and type coercion, which is the Qwen Python parser; its renderer notes a TODO to match Python `json.dumps` spacing for nested JSON. [source]
- Number: the model writes `1.0` for an integer-or-number parameter; parse returns int 1; render prints `1`. [source]
- Boolean: the model writes `True`; parse returns true; render prints `true`. [source]
- Null: `NULL` parses to nil and the renderer prints `null`. [source]
- Whitespace: a model that puts two newlines before a tool call, or no blank line, comes back with the template's own spacing. [source]
- Reasoning: leading or trailing whitespace in generated thinking is trimmed on parse and again by `TrimSpace` on render, so the re-rendered think block can be shorter than the generated one. [source]
- The Ollama Qwen3.5 parser injects a `</think>` when it sees `<tool_call>` inside thinking; the model's original bytes had no `</think>` there, so a replay that renders `</think>` after the reasoning differs from what the slot generated. [source]
- I found no test named for round trip in llama.cpp `tests/test-chat.cpp`; the one match for "round-trip" is a comment on escaped quotes in the LFM2 parser test, which tests parse only. [source]
- Whether a byte-exact round trip is a goal. Template maintainers aim at a model-faithful render (Python parity); cache-focused users need generated-bytes parity. A normalising parser and a normalising renderer each satisfy the first and break the second. [source]
- No source runs generated-text to rendered-text equality at scale on a real model. The corpus step in the proposed harness below has not been run. [source]
- Whether llama.cpp's PEG parser plus `func_args_not_string` plus the template reproduce whitespace inside argument values exactly (open in exact-tool-call-replay-for-byte-stable-agent-pre.md; still open). [source]
- Ollama's Qwen3-Coder parser removes trailing whitespace before `<tool_call>` and leading whitespace after `</tool_call>`. [source]
- Ollama's Qwen3-Coder parser strips at most one leading and one trailing newline from a parameter value, and the tests check that a newline must be the first character to be trimmed. [source]
- Ollama parses a number with no decimal part as an integer, `null` in any case as nil, and an invalid boolean as false. [source]
- The Ollama Qwen3-Coder renderer prints non-string scalars with `%v` and maps and slices as compact JSON. [source]
- The Ollama Qwen3.5 renderer applies `TrimSpace` to message content and to stored thinking before rendering them. [source]
- The Ollama Qwen3.5 parser drops whitespace after `</think>` and injects `</think>` before a `<tool_call>` found inside thinking. [source]
- llama.cpp `test_peg_parser` feeds the input one character at a time with UTF-8 safe truncation and checks the accumulated message. [source]
- llama.cpp `test_template_generation_prompt` covers content-continuation and reasoning-continuation modes. [source]
- A model that emits `1.0` for a numeric parameter is replayed as `1` by Ollama's parser and renderer, so the prefix differs from the generated tokens. [source]
- A model that emits `True` for a boolean parameter is replayed as `true` by Ollama's parser and renderer. [source]
- Neither llama.cpp's nor Ollama's test suite checks that render(parse(generated text)) equals the generated text. [source]
Children
- No children recorded.