<!-- llms-explorer concept facts · https://llms-explorer.com/tree/gemma-4-thought-stripping-in-agentic-turns/ · pack 2026-10-05 · ~739 tokens -->

# Gemma 4 thought stripping in agentic turns

> Gemma 4 guidance: before sending history for the next ordinary turn, remove the model's generated thoughts from the previous turn

Parent: [Mac local LLMs: Chat templates, reasoning and tool calling](https://llms-explorer.com/tree/mac-local-llms-chat-templates-reasoning-and-tool-calling/) · 1 facets · 10 facts · page: https://llms-explorer.com/tree/gemma-4-thought-stripping-in-agentic-turns/

## Facts

- Gemma 4 guidance: before sending history for the next ordinary turn, remove the model's generated thoughts from the previous turn — [source](https://ai.google.dev/gemma/docs/core/prompt-formatting-gemma4)
- Gemma 4 guidance: to turn thinking off mid-conversation, drop the `<|think|>` token from the system turn at the same time as stripping the previous thoughts — [source](https://ai.google.dev/gemma/docs/core/prompt-formatting-gemma4)
- Gemma 4 guidance: if one model turn contains function or tool calls, thoughts must not be removed between those calls — [source](https://ai.google.dev/gemma/docs/core/prompt-formatting-gemma4)
- Google recommends, for long-running agents, extracting and summarising earlier thoughts and feeding the summary back as ordinary text, to reduce cyclical reasoning loops caused by stripping — [source](https://ai.google.dev/gemma/docs/core/prompt-formatting-gemma4)
- Google says Gemma 4 was not trained with raw thoughts in the prompt outside the tool-call case, so no format is expected for injected thought summaries — [source](https://ai.google.dev/gemma/docs/core/prompt-formatting-gemma4)
- In Google's agentic history format the assistant message carries `tool_calls`, `tool_responses` and the final `content` together in one assistant entry, and no thought text appears in the stored history — [source](https://ai.google.dev/gemma/docs/core/prompt-formatting-gemma4)
- Google notes the 26B-A4B and 31B variants may emit a thought channel even with thinking off, and suggests adding an empty thinking token to the prompt to stabilise them — [source](https://ai.google.dev/gemma/docs/core/prompt-formatting-gemma4)
- Google suggests that when fine-tuning the 26B-A4B and 31B on data without thinking, adding the empty `<|channel>thought` block to the training prompts gives better results — [source](https://ai.google.dev/gemma/docs/core/prompt-formatting-gemma4)
- Inferred: because the previous turn's thought text is removed from the next request, a server that cached the previous turn's generated tokens (thought included) cannot reuse the KV cache past the start of that thought; the cache break is expected and is the cost of following the guidance, unlike gpt-oss within a tool loop. — source: `asserted`
- Inferred: a client that echoes `reasoning_content` back to a Gemma 4 server on an ordinary turn works against this guidance unless the template strips it; the chat template, not the client, is the safe place for the strip. — source: `asserted`
