Gemma 4 thought stripping in agentic turns
Parent: Mac local LLMs: Chat templates, reasoning and tool calling · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Gemma 4 guidance: before sending history for the next ordinary turn, remove the model's generated thoughts from the previous turn
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Gemma 4 guidance: before sending history for the next ordinary turn, remove the model's generated thoughts from the previous turn [source]
- Gemma 4 guidance: to turn thinking off mid-conversation, drop the `<|think|>` token from the system turn at the same time as stripping the previous thoughts [source]
- Gemma 4 guidance: if one model turn contains function or tool calls, thoughts must not be removed between those calls [source]
- Google recommends, for long-running agents, extracting and summarising earlier thoughts and feeding the summary back as ordinary text, to reduce cyclical reasoning loops caused by stripping [source]
- Google says Gemma 4 was not trained with raw thoughts in the prompt outside the tool-call case, so no format is expected for injected thought summaries [source]
- In Google's agentic history format the assistant message carries `tool_calls`, `tool_responses` and the final `content` together in one assistant entry, and no thought text appears in the stored history [source]
- Google notes the 26B-A4B and 31B variants may emit a thought channel even with thinking off, and suggests adding an empty thinking token to the prompt to stabilise them [source]
- Google suggests that when fine-tuning the 26B-A4B and 31B on data without thinking, adding the empty `<|channel>thought` block to the training prompts gives better results [source]
- Inferred: because the previous turn's thought text is removed from the next request, a server that cached the previous turn's generated tokens (thought included) cannot reuse the KV cache past the start of that thought; the cache break is expected and is the cost of following the guidance, unlike gpt-oss within a tool loop. [source]
- Inferred: a client that echoes `reasoning_content` back to a Gemma 4 server on an ordinary turn works against this guidance unless the template strips it; the chat template, not the client, is the safe place for the strip. [source]
Children
- No children recorded.