<!-- llms-explorer concept facts · https://llms-explorer.com/tree/parse-then-render-byte-round-trip-tests-for-reas/ · pack 2026-10-05 · ~1871 tokens -->

# Parse-then-render byte round-trip tests for reasoning turns

> Ollama's Qwen3-Coder parser trims trailing whitespace before a `<tool_call>` tag and leading whitespace after `</tool_call>`, and strips one leading and one trailing newline from every parameter value.

Parent: [Mac local LLMs: Chat templates, reasoning and tool calling](https://llms-explorer.com/tree/mac-local-llms-chat-templates-reasoning-and-tool-calling/) · 1 facets · 29 facts · page: https://llms-explorer.com/tree/parse-then-render-byte-round-trip-tests-for-reas/

## Facts

- Ollama's Qwen3-Coder parser trims trailing whitespace before a `<tool_call>` tag and leading whitespace after `</tool_call>`, and strips one leading and one trailing newline from every parameter value. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/parsers/qwen3coder.go)
- The Ollama renderer then writes a tool call as `\n<tool_call>\n<function=NAME>` plus, per argument, `\n<parameter=NAME>\n` VALUE `\n</parameter>`, so the removed newlines come back only if the model used exactly that layout. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/renderers/qwen3coder.go)
- Values are coerced by schema type on parse: `null` in any case becomes nil, boolean accepts `true` and `false` in any case, a number with no decimal part becomes an integer, and an invalid boolean becomes false. The renderer prints booleans and numbers with `%v`, strings raw, and maps and slices as compact JSON. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/parsers/qwen3coder.go)
- The Ollama Qwen3.5 renderer runs `strings.TrimSpace` on the stored thinking and on message content before writing `<think>\n` REASONING `\n</think>\n\n` CONTENT. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/renderers/qwen35.go)
- The Ollama Qwen3.5 parser eats whitespace after `</think>` and trims leading whitespace in the thinking, so generated whitespace in both places is dropped on parse. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/parsers/qwen35.go)
- llama.cpp's PEG parser tests stream the input one character at a time through `parser.parse(prefix, is_partial)`, with a UTF-8 safe truncation, and compare the accumulated message with the expectation. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tests/test-chat.cpp)
- `test_template_generation_prompt` in llama.cpp checks the generation prompt string a template adds for plain, content-continuation and reasoning-continuation cases. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tests/test-chat.cpp)
- Ollama's parser carries a comment that it follows a "reference implementation" for newline trimming and type coercion, which is the Qwen Python parser; its renderer notes a TODO to match Python `json.dumps` spacing for nested JSON. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/renderers/qwen3coder.go)
- Number: the model writes `1.0` for an integer-or-number parameter; parse returns int 1; render prints `1`. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/parsers/qwen3coder.go)
- Boolean: the model writes `True`; parse returns true; render prints `true`. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/parsers/qwen3coder.go)
- Null: `NULL` parses to nil and the renderer prints `null`. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/renderers/qwen3coder.go)
- Whitespace: a model that puts two newlines before a tool call, or no blank line, comes back with the template's own spacing. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/parsers/qwen3coder.go)
- Reasoning: leading or trailing whitespace in generated thinking is trimmed on parse and again by `TrimSpace` on render, so the re-rendered think block can be shorter than the generated one. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/renderers/qwen35.go)
- The Ollama Qwen3.5 parser injects a `</think>` when it sees `<tool_call>` inside thinking; the model's original bytes had no `</think>` there, so a replay that renders `</think>` after the reasoning differs from what the slot generated. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/parsers/qwen35.go)
- I found no test named for round trip in llama.cpp `tests/test-chat.cpp`; the one match for "round-trip" is a comment on escaped quotes in the LFM2 parser test, which tests parse only. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tests/test-chat.cpp)
- Whether a byte-exact round trip is a goal. Template maintainers aim at a model-faithful render (Python parity); cache-focused users need generated-bytes parity. A normalising parser and a normalising renderer each satisfy the first and break the second. — source: `asserted`
- No source runs generated-text to rendered-text equality at scale on a real model. The corpus step in the proposed harness below has not been run. — source: `asserted`
- Whether llama.cpp's PEG parser plus `func_args_not_string` plus the template reproduce whitespace inside argument values exactly (open in exact-tool-call-replay-for-byte-stable-agent-pre.md; still open). — source: `asserted`
- Ollama's Qwen3-Coder parser removes trailing whitespace before `<tool_call>` and leading whitespace after `</tool_call>`. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/parsers/qwen3coder.go)
- Ollama's Qwen3-Coder parser strips at most one leading and one trailing newline from a parameter value, and the tests check that a newline must be the first character to be trimmed. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/parsers/qwen3coder_test.go)
- Ollama parses a number with no decimal part as an integer, `null` in any case as nil, and an invalid boolean as false. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/parsers/qwen3coder.go)
- The Ollama Qwen3-Coder renderer prints non-string scalars with `%v` and maps and slices as compact JSON. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/renderers/qwen3coder.go)
- The Ollama Qwen3.5 renderer applies `TrimSpace` to message content and to stored thinking before rendering them. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/renderers/qwen35.go)
- The Ollama Qwen3.5 parser drops whitespace after `</think>` and injects `</think>` before a `<tool_call>` found inside thinking. — [source](https://raw.githubusercontent.com/ollama/ollama/main/model/parsers/qwen35.go)
- llama.cpp `test_peg_parser` feeds the input one character at a time with UTF-8 safe truncation and checks the accumulated message. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tests/test-chat.cpp)
- llama.cpp `test_template_generation_prompt` covers content-continuation and reasoning-continuation modes. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tests/test-chat.cpp)
- A model that emits `1.0` for a numeric parameter is replayed as `1` by Ollama's parser and renderer, so the prefix differs from the generated tokens. — source: `asserted`
- A model that emits `True` for a boolean parameter is replayed as `true` by Ollama's parser and renderer. — source: `asserted`
- Neither llama.cpp's nor Ollama's test suite checks that render(parse(generated text)) equals the generated text. — source: `asserted`
