Responses API client-executed tool_search support in local servers
Parent: Mac local LLMs: Agent clients, context and compaction · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
OpenAI's tool-search guide defines two modes: hosted, where OpenAI searches the deferred tools declared in the request and returns the loaded subset in the same response, and client-executed, where the model emits a `tool_search_call`, the application does the lookup and returns a `tool_search_ou...
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- OpenAI's tool-search guide defines two modes: hosted, where OpenAI searches the deferred tools declared in the request and returns the loaded subset in the same response, and client-executed, where the model emits a `tool_search_call`, the application does the lookup and returns a `tool_search_output`. [source]
- Client mode is configured by adding a `tool_search` tool with `execution: "client"` and a parameters schema for the arguments the application expects. [source]
- The client returns `{"type": "tool_search_output", "execution": "client", "call_id": <the call's id>, "status": "completed", "tools": [...]}` in the next request's `input`, after the earlier output items. [source]
- In hosted mode the `tool_search_call` and `tool_search_output` items carry `execution: "server"` and `call_id: null`. [source]
- `tool_search_output.tools` lists the tools the model can call on later turns; tools not in the array are not loaded, and the client need not load the same tool again on later turns. [source]
- Client mode also permits returning tools that were absent from the original `tools` list; OpenAI calls this an advanced workflow and says to validate the returned schemas carefully. [source]
- An `additional_tools` input item makes its tools available only from the point where the item appears, so a client that replays items by hand must keep the item's position. [source]
- Loaded tools are placed at the end of the model's context window in both modes so the cached prefix survives. [source]
- A LiteLLM Responses-to-chat translation layer rejected Codex's `type: "tool_search"` entry with `400: Field required ... tools[13].function`, which shows a chat-translating router cannot carry the item type unchanged. [source]
- OpenAI's guide limits `tool_search` in the Responses API to `gpt-5.4` and later, so a local model gets it only if a server or proxy emulates the item types. [source]
Children
- No children recorded.