<!-- llms-explorer concept facts · https://llms-explorer.com/tree/gpt-oss-harmony-reasoning-replay-across-tool-cal/ · pack 2026-10-05 · ~1186 tokens -->

# gpt-oss Harmony reasoning replay across tool calls

> Harmony rule 1: chain of thought is issued to the `analysis` channel

Parent: [Mac local LLMs: Chat templates, reasoning and tool calling](https://llms-explorer.com/tree/mac-local-llms-chat-templates-reasoning-and-tool-calling/) · 1 facets · 17 facts · page: https://llms-explorer.com/tree/gpt-oss-harmony-reasoning-replay-across-tool-cal/

## Facts

- Harmony rule 1: chain of thought is issued to the `analysis` channel — [source](https://cookbook.openai.com/articles/gpt-oss/handle-raw-cot)
- Harmony rule 2: once a later sampling turn follows a message on the `final` channel, all earlier `analysis` messages should be dropped, while function calls on the `commentary` channel can remain — [source](https://cookbook.openai.com/articles/gpt-oss/handle-raw-cot)
- Harmony rule 3: if the last assistant message was a tool call of any type, the analysis messages back to the previous `final` message must be kept on the next sampling until a `final` message is issued — [source](https://cookbook.openai.com/articles/gpt-oss/handle-raw-cot)
- The Harmony guide's worked example rewrites the assistant's closing `<|return|>` as `<|end|>` when the finished answer is replayed in the next prompt, and omits its `analysis` message — [source](https://cookbook.openai.com/articles/openai-harmony)
- Harmony states that gpt-oss can call tools as part of its chain of thought, and that this is the single exception to dropping prior CoT — [source](https://cookbook.openai.com/articles/openai-harmony)
- Harmony stop tokens are `<|return|>` (id 200002, final answer done) and `<|call|>` (id 200012, tool call requested); `<|constrain|>` is id 200003 and `<|channel|>` is id 200005 — [source](https://cookbook.openai.com/articles/openai-harmony)
- Harmony routes function tool calls to `commentary` and built-in tools normally to `analysis`, with occasional exceptions, so a parser must not assume a fixed channel per tool kind — [source](https://cookbook.openai.com/articles/openai-harmony)
- Harmony tells implementers that analysis-channel text does not meet the safety standard of `final` text and must not be shown to end users — [source](https://cookbook.openai.com/articles/openai-harmony)
- OpenAI's raw-CoT guide says the CoT is also crucial for tool-calling performance, not only for safety research — [source](https://cookbook.openai.com/articles/gpt-oss/handle-raw-cot)
- Failure signature when replay is missing or the field name mismatches: the model needs the CoT from the previous tool call but the client does not send it; aldehir (llama.cpp collaborator, 2025-08-18) called this an open issue across all clients and inference servers for gpt-oss — [source](https://github.com/ggml-org/llama.cpp/issues/15396)
- llama-server only honours `chat_template_kwargs.reasoning_effort` from the request or the `--chat-template-kwargs` default; a client that never sets it gets the server default, and the guide says client requests override server defaults — [source](https://github.com/ggml-org/llama.cpp/issues/15396)
- The llama.cpp gpt-oss guide recommends `--temp 1.0 --top-p 1.0` (OpenAI's recommended sampling) and `--reasoning-format auto`; `none` was no longer recommended as of 2025-08-19 — [source](https://github.com/ggml-org/llama.cpp/issues/15396)
- Cline and Roo Code did not use native tool calls in 2025-08, and gpt-oss insisted on native Harmony calls; a workaround grammar `root ::= analysis? start final .+` passed via `--grammar-file` plus the system line 'Valid channels: analysis, final.' blocked the commentary channel, and the author said the 20B model failed without it — [source](https://github.com/ggml-org/llama.cpp/issues/15396)
- The same guide thread says the proper fix for the non-native-tool clients is native tool calling in the client (a Roo Code PR was open), and calls the grammar a hack providers will not adopt — [source](https://github.com/ggml-org/llama.cpp/issues/15396)
- A llama-server run on 2025-08 failed OpenAI's gpt-oss compatibility test (30 of 30 cases) only because the test read `reasoning` while llama-server returned `reasoning_content`; LM Studio passed — [source](https://github.com/ggml-org/llama.cpp/discussions/15362)
- In the Responses API OpenAI added a `content` array of `reasoning_text` items on `reasoning` items plus `response.reasoning_text.delta` and `.done` events, so raw CoT and the user-visible `summary` travel together — [source](https://cookbook.openai.com/articles/gpt-oss/handle-raw-cot)
- Inferred: a stateless Chat Completions client that rebuilds history from `content` and `tool_calls` only will silently drop the analysis text and degrade multi-step gpt-oss tool use without any error. — source: `asserted`
