Unsloth Studio as an Anthropic-compatible endpoint for Claude Code
Parent: Mac local LLMs: Runtime selection and frontends · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Unsloth's API page states Unsloth "speaks two dialects on the same port": Anthropic-compatible `/v1/messages` (Claude Code, OpenClaw, Anthropic SDK) and OpenAI-compatible `/v1/chat/completions` and `/v1/responses`.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Unsloth's API page states Unsloth "speaks two dialects on the same port": Anthropic-compatible `/v1/messages` (Claude Code, OpenClaw, Anthropic SDK) and OpenAI-compatible `/v1/chat/completions` and `/v1/responses`. [source]
- Unsloth's Claude Code guide documents `unsloth start claude --model <repo>:<quant>` with sampling flags and `--reasoning-effort medium`, and says that with no flags Unsloth picks recommended settings including context length and temperature. [source]
- Manual setup is `ANTHROPIC_BASE_URL=http://localhost:8888`, `ANTHROPIC_AUTH_TOKEN` set to the `sk-unsloth-...` key, an empty `ANTHROPIC_API_KEY` so Claude Code does not ask for a cloud key, and `ANTHROPIC_MODEL` set to the exact ID returned by `GET /v1/models`. [source]
- Unsloth's guide says Claude Code "recently prepends" an attribution header (`x-anthropic-billing-header: cc_version=...; cch=...;`) whose value changes on every request, making inference "90% slower with local models", and gives `CLAUDE_CODE_ATTRIBUTION_HEADER=0` (with `CLAUDE_CODE_ENABLE_TELEMETRY=0`) to turn it off. [source]
- The guide says older Claude Code builds ignored the shell variable `CLAUDE_CODE_ATTRIBUTION_HEADER=0`, so the `--settings` form or the settings-file `env` block is the reliable choice. [source]
- llama.cpp's Anthropic converter rewrites the five-character `cch=` value of a leading `x-anthropic-billing-header:` system prompt to `fffff`, so the prefix cache survives; the change is llama.cpp PR 21793, commit c807c6e3 dated 2026-04-23. [source]
- The guide's prompt-shrinking flags are `claude --bare --exclude-dynamic-system-prompt-sections`: `--bare` skips auto-discovery of hooks, skills, plugins, MCP servers and CLAUDE.md, and the second flag moves per-machine sections out of the prompt prefix. [source]
- The guide says to start the server for a coding agent with `unsloth run --model ... --disable-tools --reasoning off -p 8888`, because by default Studio "runs its own server-side tools, which swallows the agent's tool calls, so Claude Code answers but never edits files"; `--disable-tools` switches to passthrough. [source]
- Server-side tool policy depends on the bind address: tools are on by default on `127.0.0.1` and off by default on `0.0.0.0` or any non-loopback address; `--enable-tools` on `0.0.0.0` shows a y/N prompt (skip with `--yes`) and the resolved policy cannot be overridden per request. [source]
- Unsloth maps Anthropic `tool_choice` as `auto` to OpenAI `auto`, `any` to `required`, `{type: tool, name}` to `{type: function, function: {name}}`, and `none` to `none`. [source]
- llama-server's own Anthropic converter maps `tool_choice` `auto` to `auto` and both `any` and `tool` to `required`, dropping the tool name, and does not map `none`. [source]
- Unsloth says server-sent events do not survive the `unsloth studio --secure` Cloudflare quick tunnel and clients must set `stream: false` over it. [source]
- Unsloth's troubleshooting says `Unable to connect to API (ConnectionRefused)` after returning to the cloud is fixed by `unset ANTHROPIC_BASE_URL`, and a `stream: true` request that returns one JSON blob means the client is buffering or the path is wrong. [source]
- In April 2026 users reported Studio's inference API as experimental, with `/v1/chat/completions` defaulting to streaming when `stream` is omitted (issue 5047) and tool calling broken through the proxy (issue 4999). [source]
- Unsloth PR 5122 extended client-side tool pass-through (from PR 5099) to `/v1/responses` so Codex CLI gets structured `function_call` items instead of raw tool-call tokens in `content`. [source]
- Inferred: a Claude Code user on Studio who forgets `--disable-tools` on localhost gets a working chat but no file edits, which looks like a model failure and not a configuration one. [source]
Children
- No children recorded.