OpenAI gpt-oss compatibility test script for verifying server implementations
Parent: Mac local LLMs: Chat templates, reasoning and tool calling · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Files: `index.ts`, `providers.ts`, `runCase.ts`, `tools.ts`, `analysis.ts`, `cases.jsonl`, `package.json`.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Files: `index.ts`, `providers.ts`, `runCase.ts`, `tools.ts`, `analysis.ts`, `cases.jsonl`, `package.json`. [source]
- Steps: `npm install`; add a provider entry in `providers.ts` (`chat` for Chat Completions, `responses` for the Responses API); smoke test with `npm start -- --provider <name> -n 1 -k 1`; full run `npm start -- --provider <name> -k 5` (each case 5 times for consistency). [source]
- Aug 2025: llama.cpp discussion 15362 used the script and showed 30 of 30 failures on llama-server from the `reasoning` field name, which led to the `--reasoning-format` discussion. [source]
- The script fails when the API shape differs; chat-API streaming events are not tested; a call that passes schema validation with wrong input still passes. [source]
- A collaborator argued some "failures" on gpt-oss-20b are the model adding optional parameters the test does not expect, which is within the script's margin of error. [source]
- Whether current llama-server builds pass the script; no 2026 run found. [source]
- The gpt-oss compatibility test is a Node/TypeScript project, not a Python script, and needs a `providers.ts` entry for the server under test. [source]
- `-n` limits the number of cases and `-k` sets repeats per case in `npm start`. [source]
- The script covers both Chat Completions and Responses API providers via the provider type. [source]
- The repo also ships an example Responses API server that does not implement every event but should cover basic use cases. [source]
- Optional-parameter differences in gpt-oss output are not treated as failures by a llama.cpp collaborator. [source]
- The 2025-08 failure count (30 of 30) is already in gpt-oss-harmony-reasoning-replay-across-tool-cal.md; this file adds only the script mechanics. [source]
Children
- No children recorded.