Reasoning-block trailing newline parser bug causing agent loops
Parent: Mac local LLMs: Chat templates, reasoning and tool calling · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
The loop was reported as llama.cpp issue 24181, titled "Eval bug: Step 3.7 Flash gets stuck in reasoning trying to make tool calls (autoparser)", renamed on 2026-06-05 from "Bug: Step 3.7 Flash enters infinite reasoning loops".
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- The loop was reported as llama.cpp issue 24181, titled "Eval bug: Step 3.7 Flash gets stuck in reasoning trying to make tool calls (autoparser)", renamed on 2026-06-05 from "Bug: Step 3.7 Flash enters infinite reasoning loops". [source]
- A Hacker News commenter who debugged it says the parser captured an extra `\n` as part of the reasoning block, it appeared only in longer multi-turn agentic sessions, and the extra linefeed steered the model into reasoning self-corrections that worsened with session length. [source]
- A second commenter's reading of the issue thread: a `\n` slipped through from the last line of the reasoning trace, producing a blank line before `</think>` that occasionally triggered an "Actually ..." digression which then looped. [source]
- The first commenter explains the loop as incremental in-context imitation: a model trained to end reasoning with one linefeed then `</think>` sees two consecutive linefeeds in its replayed history, adds one itself, and each digression raises the chance of the next. [source]
- The same commenter says the autoparser "was (and still is)" parsing a trailing linefeed as part of a reasoning block and that a complete fix needs both a parser-definition fix and whitespace trimming in the encoding step, since the latter also covers clients that add the whitespace deliberately. [source]
- He states the maintainer fixed this specific issue by trimming the extra linefeed before passing the text to the template. [source]
- He describes the autoparser's goal as inferring delimiter tokens from the chat template so new model parsers are easier to write, and notes that llama.cpp's API returns pre-parsed data so clients need not parse thinking, text or tool blocks. [source]
- In issue 24181 the autoparser found raw reasoning delimiters with surrounding whitespace (start `<think>\n`, end `\n</think>\n`) while `common_chat_params` exposed only the trimmed tags. [source]
- A patch posted in issue 24181 opts Step 3.7 Flash out of the autoparser with a dedicated parser: the generation prompt is `<|im_start|>assistant\n<think>\n` plus reasoning, `\n</think>\n` and content; the open delimiter is `<think>\n` or `<think>` with an optional newline; the close is an optional `\n` before `</think>` and an optional `\n` after; and reasoning text is read up to `\n</think>` or `</think>`. [source]
- The same patch uses a lazy grammar triggered by the word `<tool_call>` when tools are present and tool choice is auto. [source]
- Inferred: the delimiter's newline must belong to the delimiter (consumed by the parser) and not to the reasoning text, otherwise every round trip through a non-trimming template adds a line feed. [source]
- Inferred: a server-side regression test should replay a reasoning turn through parse then template and assert the rendered bytes equal the generated bytes. [source]
Children
- No children recorded.