gpt-oss Harmony reasoning replay across tool calls
Parent: Mac local LLMs: Chat templates, reasoning and tool calling · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Harmony rule 1: chain of thought is issued to the `analysis` channel
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Harmony rule 1: chain of thought is issued to the `analysis` channel [source]
- Harmony rule 2: once a later sampling turn follows a message on the `final` channel, all earlier `analysis` messages should be dropped, while function calls on the `commentary` channel can remain [source]
- Harmony rule 3: if the last assistant message was a tool call of any type, the analysis messages back to the previous `final` message must be kept on the next sampling until a `final` message is issued [source]
- The Harmony guide's worked example rewrites the assistant's closing `<|return|>` as `<|end|>` when the finished answer is replayed in the next prompt, and omits its `analysis` message [source]
- Harmony states that gpt-oss can call tools as part of its chain of thought, and that this is the single exception to dropping prior CoT [source]
- Harmony stop tokens are `<|return|>` (id 200002, final answer done) and `<|call|>` (id 200012, tool call requested); `<|constrain|>` is id 200003 and `<|channel|>` is id 200005 [source]
- Harmony routes function tool calls to `commentary` and built-in tools normally to `analysis`, with occasional exceptions, so a parser must not assume a fixed channel per tool kind [source]
- Harmony tells implementers that analysis-channel text does not meet the safety standard of `final` text and must not be shown to end users [source]
- OpenAI's raw-CoT guide says the CoT is also crucial for tool-calling performance, not only for safety research [source]
- Failure signature when replay is missing or the field name mismatches: the model needs the CoT from the previous tool call but the client does not send it; aldehir (llama.cpp collaborator, 2025-08-18) called this an open issue across all clients and inference servers for gpt-oss [source]
- llama-server only honours `chat_template_kwargs.reasoning_effort` from the request or the `--chat-template-kwargs` default; a client that never sets it gets the server default, and the guide says client requests override server defaults [source]
- The llama.cpp gpt-oss guide recommends `--temp 1.0 --top-p 1.0` (OpenAI's recommended sampling) and `--reasoning-format auto`; `none` was no longer recommended as of 2025-08-19 [source]
- Cline and Roo Code did not use native tool calls in 2025-08, and gpt-oss insisted on native Harmony calls; a workaround grammar `root ::= analysis? start final .+` passed via `--grammar-file` plus the system line 'Valid channels: analysis, final.' blocked the commentary channel, and the author said the 20B model failed without it [source]
- The same guide thread says the proper fix for the non-native-tool clients is native tool calling in the client (a Roo Code PR was open), and calls the grammar a hack providers will not adopt [source]
- A llama-server run on 2025-08 failed OpenAI's gpt-oss compatibility test (30 of 30 cases) only because the test read `reasoning` while llama-server returned `reasoning_content`; LM Studio passed [source]
- In the Responses API OpenAI added a `content` array of `reasoning_text` items on `reasoning` items plus `response.reasoning_text.delta` and `.done` events, so raw CoT and the user-visible `summary` travel together [source]
- Inferred: a stateless Chat Completions client that rebuilds history from `content` and `tool_calls` only will silently drop the analysis text and degrade multi-step gpt-oss tool use without any error. [source]
Children
- No children recorded.