Codex Responses API compatibility on local servers
Parent: Mac local LLMs: Chat templates, reasoning and tool calling · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Codex's `model_providers.<id>.wire_api` accepts only `responses`, which is also the default when omitted.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Codex's `model_providers.<id>.wire_api` accepts only `responses`, which is also the default when omitted. [source]
- Codex reserves the built-in provider IDs `openai`, `ollama` and `lmstudio`; `--oss` runs against LM Studio or Ollama and `oss_provider` sets the default. [source]
- Codex provider keys include `stream_idle_timeout_ms` (default 300000), `supports_websockets`, `query_params` and `env_key`. [source]
- In an April 2026 test with Codex CLI v0.123.0, `/v1/responses` worked for direct non-agent requests on llama.cpp but a real `codex exec` failed with HTTP 400 `'type' of tool must be 'function'`. [source]
- The captured Codex `tools[]` held `function`, `web_search`, `image_generation` and `namespace` entries, and a small local proxy that dropped the non-function entries let Codex finish. [source]
- On 2026-06-18, with codex-cli 0.141.0, the same capture still held function, several namespace, web_search and image_generation tools. [source]
- An Unsloth maintainer stated Codex talks only to `/responses` and has dropped `/v1/chat/completions`, and that a GGUF served through Unsloth Studio exposes a tested `/responses` endpoint. [source]
- llama.cpp commit 91e84fed (2026-05-15), "Support for Codex CLI by skipping unsupported Responses tools (#23041)", touches `tools/server/server-chat.cpp`. [source]
- On master, the Responses converter skips any tool whose `type` is not `function` with the warning "unsupported Responses tool type '%s' skipped", and removes the `tools` key if none remain. [source]
- The converter sets `strict` to true on a function tool that lacks it. [source]
- The converter turns `instructions` into a leading system message and keeps `user`, `system` and `developer` input messages with their original role. [source]
- The converter rejects `previous_response_id` with "llama.cpp does not support 'previous_response_id'." and rejects `input_file` parts. [source]
- A `reasoning` input item is converted to `reasoning_content` from `content[0].text`, and the converter throws if `content` is missing, empty or its first part has no text. [source]
- `max_output_tokens` becomes `max_tokens`, and of the `reasoning` object only `effort` is mapped. [source]
- llama-server also serves `POST /v1/responses/input_tokens` for token counting and states `/v1/responses` works by converting the request into a Chat Completions request. [source]
- llama.cpp PR 18486 (2025-12-30) emits Responses SSE and flags a caveat: all `response.output_item.done` events come at the end, which Codex does not mind because it does not check event order. [source]
- llama.cpp commit 4098fdc9 (2026-09-22) added `input_image` support inside `function_call_output`. [source]
- A April 2026 Gemma 4 test needed `web_search = "disabled"` in the Codex profile because Codex sent a `web_search_preview` tool type that llama.cpp (build 8680) rejected. [source]
- The same test advises `stream_idle_timeout_ms` of at least 1,800,000 because one tool-call cycle took 1 minute 39 seconds on a 24 GB M4 Pro. [source]
- Unsloth Studio PR 5122 translates flat Responses function tools to nested Chat Completions tools, maps `function_call_output` to `role: tool` and `function_call` to assistant `tool_calls`, and before it `/v1/responses` silently dropped `tools`. [source]
- PR 5122 merges every `instructions` and `developer` or `system` fragment into one leading `system` message because Codex sends both and strict templates raise "System message must be at the beginning." on the second. [source]
- PR 5122 fixes HTTP 422 on turn two onward by accepting `output_text` assistant replay parts, `reasoning` items and the `phase` field Codex attaches to assistant messages, and rejects non-GGUF streaming with a typed 400. [source]
- Unsloth's Codex guide, per the June 2026 re-check, still shows raw llama.cpp on port 8001 with `wire_api = "responses"` and omits `--ctx-size`. [source]
- A September 2026 how-to shows `codex --oss` posting to LM Studio at `http://localhost:1234/v1/responses` and an Ollama profile with `wire_api = "responses"`. [source]
- Inferred: because a `namespace` tool is skipped whole, any functions Codex nests in it are not offered to a local model behind llama-server. [source]
Children
- No children recorded.