<!-- llms-explorer concept facts · https://llms-explorer.com/tree/codex-responses-api-compatibility-on-local-serve/ · pack 2026-10-05 · ~1603 tokens -->

# Codex Responses API compatibility on local servers

> Codex's `model_providers.<id>.wire_api` accepts only `responses`, which is also the default when omitted.

Parent: [Mac local LLMs: Chat templates, reasoning and tool calling](https://llms-explorer.com/tree/mac-local-llms-chat-templates-reasoning-and-tool-calling/) · 1 facets · 25 facts · page: https://llms-explorer.com/tree/codex-responses-api-compatibility-on-local-serve/

## Facts

- Codex's `model_providers.<id>.wire_api` accepts only `responses`, which is also the default when omitted. — [source](https://developers.openai.com/codex/config-reference)
- Codex reserves the built-in provider IDs `openai`, `ollama` and `lmstudio`; `--oss` runs against LM Studio or Ollama and `oss_provider` sets the default. — [source](https://developers.openai.com/codex/config-advanced)
- Codex provider keys include `stream_idle_timeout_ms` (default 300000), `supports_websockets`, `query_params` and `env_key`. — [source](https://developers.openai.com/codex/config-reference)
- In an April 2026 test with Codex CLI v0.123.0, `/v1/responses` worked for direct non-agent requests on llama.cpp but a real `codex exec` failed with HTTP 400 `'type' of tool must be 'function'`. — [source](https://github.com/unslothai/unsloth/issues/5141)
- The captured Codex `tools[]` held `function`, `web_search`, `image_generation` and `namespace` entries, and a small local proxy that dropped the non-function entries let Codex finish. — [source](https://github.com/unslothai/unsloth/issues/5141)
- On 2026-06-18, with codex-cli 0.141.0, the same capture still held function, several namespace, web_search and image_generation tools. — [source](https://github.com/unslothai/unsloth/issues/5141)
- An Unsloth maintainer stated Codex talks only to `/responses` and has dropped `/v1/chat/completions`, and that a GGUF served through Unsloth Studio exposes a tested `/responses` endpoint. — [source](https://github.com/unslothai/unsloth/issues/5141)
- llama.cpp commit 91e84fed (2026-05-15), "Support for Codex CLI by skipping unsupported Responses tools (#23041)", touches `tools/server/server-chat.cpp`. — [source](https://api.github.com/repos/ggml-org/llama.cpp/commits?path=tools/server/server-chat.cpp&per_page=30)
- On master, the Responses converter skips any tool whose `type` is not `function` with the warning "unsupported Responses tool type '%s' skipped", and removes the `tools` key if none remain. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-chat.cpp)
- The converter sets `strict` to true on a function tool that lacks it. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-chat.cpp)
- The converter turns `instructions` into a leading system message and keeps `user`, `system` and `developer` input messages with their original role. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-chat.cpp)
- The converter rejects `previous_response_id` with "llama.cpp does not support 'previous_response_id'." and rejects `input_file` parts. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-chat.cpp)
- A `reasoning` input item is converted to `reasoning_content` from `content[0].text`, and the converter throws if `content` is missing, empty or its first part has no text. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-chat.cpp)
- `max_output_tokens` becomes `max_tokens`, and of the `reasoning` object only `effort` is mapped. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-chat.cpp)
- llama-server also serves `POST /v1/responses/input_tokens` for token counting and states `/v1/responses` works by converting the request into a Chat Completions request. — [source](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)
- llama.cpp PR 18486 (2025-12-30) emits Responses SSE and flags a caveat: all `response.output_item.done` events come at the end, which Codex does not mind because it does not check event order. — [source](https://github.com/ggml-org/llama.cpp/pull/18486)
- llama.cpp commit 4098fdc9 (2026-09-22) added `input_image` support inside `function_call_output`. — [source](https://api.github.com/repos/ggml-org/llama.cpp/commits?path=tools/server/server-chat.cpp&per_page=30)
- A April 2026 Gemma 4 test needed `web_search = "disabled"` in the Codex profile because Codex sent a `web_search_preview` tool type that llama.cpp (build 8680) rejected. — [source](https://medium.com/google-cloud/i-ran-gemma-4-as-a-local-model-in-codex-cli-7fda754dc0d4)
- The same test advises `stream_idle_timeout_ms` of at least 1,800,000 because one tool-call cycle took 1 minute 39 seconds on a 24 GB M4 Pro. — [source](https://medium.com/google-cloud/i-ran-gemma-4-as-a-local-model-in-codex-cli-7fda754dc0d4)
- Unsloth Studio PR 5122 translates flat Responses function tools to nested Chat Completions tools, maps `function_call_output` to `role: tool` and `function_call` to assistant `tool_calls`, and before it `/v1/responses` silently dropped `tools`. — [source](https://github.com/unslothai/unsloth/pull/5122)
- PR 5122 merges every `instructions` and `developer` or `system` fragment into one leading `system` message because Codex sends both and strict templates raise "System message must be at the beginning." on the second. — [source](https://github.com/unslothai/unsloth/pull/5122)
- PR 5122 fixes HTTP 422 on turn two onward by accepting `output_text` assistant replay parts, `reasoning` items and the `phase` field Codex attaches to assistant messages, and rejects non-GGUF streaming with a typed 400. — [source](https://github.com/unslothai/unsloth/pull/5122)
- Unsloth's Codex guide, per the June 2026 re-check, still shows raw llama.cpp on port 8001 with `wire_api = "responses"` and omits `--ctx-size`. — [source](https://github.com/unslothai/unsloth/issues/5141)
- A September 2026 how-to shows `codex --oss` posting to LM Studio at `http://localhost:1234/v1/responses` and an Ollama profile with `wire_api = "responses"`. — [source](https://michaelwolfinger.com/blog/2026/codex-cli-local-model-lm-studio-ollama/)
- Inferred: because a `namespace` tool is skipped whole, any functions Codex nests in it are not offered to a local model behind llama-server. — source: `asserted`
