<!-- llms-explorer concept facts · https://llms-explorer.com/tree/responses-api-client-executed-tool-search-suppor/ · pack 2026-10-05 · ~801 tokens -->

# Responses API client-executed tool_search support in local servers

> OpenAI's tool-search guide defines two modes: hosted, where OpenAI searches the deferred tools declared in the request and returns the loaded subset in the same response, and client-executed, where the model emits a `tool_search_call`, the application does the lookup and returns a `tool_search_ou...

Parent: [Mac local LLMs: Agent clients, context and compaction](https://llms-explorer.com/tree/mac-local-llms-agent-clients-context-and-compaction/) · 1 facets · 10 facts · page: https://llms-explorer.com/tree/responses-api-client-executed-tool-search-suppor/

## Facts

- OpenAI's tool-search guide defines two modes: hosted, where OpenAI searches the deferred tools declared in the request and returns the loaded subset in the same response, and client-executed, where the model emits a `tool_search_call`, the application does the lookup and returns a `tool_search_output`. — [source](https://developers.openai.com/api/docs/guides/tools-tool-search)
- Client mode is configured by adding a `tool_search` tool with `execution: "client"` and a parameters schema for the arguments the application expects. — [source](https://developers.openai.com/api/docs/guides/tools-tool-search)
- The client returns `{"type": "tool_search_output", "execution": "client", "call_id": <the call's id>, "status": "completed", "tools": [...]}` in the next request's `input`, after the earlier output items. — [source](https://developers.openai.com/api/docs/guides/tools-tool-search)
- In hosted mode the `tool_search_call` and `tool_search_output` items carry `execution: "server"` and `call_id: null`. — [source](https://developers.openai.com/api/docs/guides/tools-tool-search)
- `tool_search_output.tools` lists the tools the model can call on later turns; tools not in the array are not loaded, and the client need not load the same tool again on later turns. — [source](https://developers.openai.com/api/docs/guides/tools-tool-search)
- Client mode also permits returning tools that were absent from the original `tools` list; OpenAI calls this an advanced workflow and says to validate the returned schemas carefully. — [source](https://developers.openai.com/api/docs/guides/tools-tool-search)
- An `additional_tools` input item makes its tools available only from the point where the item appears, so a client that replays items by hand must keep the item's position. — [source](https://developers.openai.com/api/docs/guides/tools-tool-search)
- Loaded tools are placed at the end of the model's context window in both modes so the cached prefix survives. — [source](https://developers.openai.com/api/docs/guides/tools-tool-search)
- A LiteLLM Responses-to-chat translation layer rejected Codex's `type: "tool_search"` entry with `400: Field required ... tools[13].function`, which shows a chat-translating router cannot carry the item type unchanged. — [source](https://github.com/openai/codex/issues/36382)
- OpenAI's guide limits `tool_search` in the Responses API to `gpt-5.4` and later, so a local model gets it only if a server or proxy emulates the item types. — [source](https://developers.openai.com/api/docs/guides/tools-tool-search)
