<!-- llms-explorer concept facts · https://llms-explorer.com/tree/openai-gpt-oss-compatibility-test-script-for-ver/ · pack 2026-10-05 · ~719 tokens -->

# OpenAI gpt-oss compatibility test script for verifying server implementations

> Files: `index.ts`, `providers.ts`, `runCase.ts`, `tools.ts`, `analysis.ts`, `cases.jsonl`, `package.json`.

Parent: [Mac local LLMs: Chat templates, reasoning and tool calling](https://llms-explorer.com/tree/mac-local-llms-chat-templates-reasoning-and-tool-calling/) · 1 facets · 12 facts · page: https://llms-explorer.com/tree/openai-gpt-oss-compatibility-test-script-for-ver/

## Facts

- Files: `index.ts`, `providers.ts`, `runCase.ts`, `tools.ts`, `analysis.ts`, `cases.jsonl`, `package.json`. — [source](https://github.com/openai/gpt-oss/tree/main/compatibility-test)
- Steps: `npm install`; add a provider entry in `providers.ts` (`chat` for Chat Completions, `responses` for the Responses API); smoke test with `npm start -- --provider <name> -n 1 -k 1`; full run `npm start -- --provider <name> -k 5` (each case 5 times for consistency). — [source](https://github.com/openai/gpt-oss/tree/main/compatibility-test)
- Aug 2025: llama.cpp discussion 15362 used the script and showed 30 of 30 failures on llama-server from the `reasoning` field name, which led to the `--reasoning-format` discussion. — [source](https://github.com/ggml-org/llama.cpp/discussions/15362)
- The script fails when the API shape differs; chat-API streaming events are not tested; a call that passes schema validation with wrong input still passes. — [source](https://github.com/openai/gpt-oss/tree/main/compatibility-test)
- A collaborator argued some "failures" on gpt-oss-20b are the model adding optional parameters the test does not expect, which is within the script's margin of error. — [source](https://github.com/ggml-org/llama.cpp/discussions/15362)
- Whether current llama-server builds pass the script; no 2026 run found. — source: `asserted`
- The gpt-oss compatibility test is a Node/TypeScript project, not a Python script, and needs a `providers.ts` entry for the server under test. — [source](https://github.com/openai/gpt-oss/tree/main/compatibility-test)
- `-n` limits the number of cases and `-k` sets repeats per case in `npm start`. — [source](https://github.com/openai/gpt-oss/tree/main/compatibility-test)
- The script covers both Chat Completions and Responses API providers via the provider type. — [source](https://github.com/openai/gpt-oss/tree/main/compatibility-test)
- The repo also ships an example Responses API server that does not implement every event but should cover basic use cases. — [source](https://github.com/openai/gpt-oss)
- Optional-parameter differences in gpt-oss output are not treated as failures by a llama.cpp collaborator. — [source](https://github.com/ggml-org/llama.cpp/discussions/15362)
- The 2025-08 failure count (30 of 30) is already in gpt-oss-harmony-reasoning-replay-across-tool-cal.md; this file adds only the script mechanics. — source: `asserted`
