<!-- llms-explorer concept facts · https://llms-explorer.com/tree/unsloth-studio-as-an-anthropic-compatible-endpoi/ · pack 2026-10-05 · ~1282 tokens -->

# Unsloth Studio as an Anthropic-compatible endpoint for Claude Code

> Unsloth's API page states Unsloth "speaks two dialects on the same port": Anthropic-compatible `/v1/messages` (Claude Code, OpenClaw, Anthropic SDK) and OpenAI-compatible `/v1/chat/completions` and `/v1/responses`.

Parent: [Mac local LLMs: Runtime selection and frontends](https://llms-explorer.com/tree/mac-local-llms-runtime-selection-and-frontends/) · 1 facets · 16 facts · page: https://llms-explorer.com/tree/unsloth-studio-as-an-anthropic-compatible-endpoi/

## Facts

- Unsloth's API page states Unsloth "speaks two dialects on the same port": Anthropic-compatible `/v1/messages` (Claude Code, OpenClaw, Anthropic SDK) and OpenAI-compatible `/v1/chat/completions` and `/v1/responses`. — [source](https://unsloth.ai/docs/basics/api)
- Unsloth's Claude Code guide documents `unsloth start claude --model <repo>:<quant>` with sampling flags and `--reasoning-effort medium`, and says that with no flags Unsloth picks recommended settings including context length and temperature. — [source](https://unsloth.ai/docs/basics/claude-code)
- Manual setup is `ANTHROPIC_BASE_URL=http://localhost:8888`, `ANTHROPIC_AUTH_TOKEN` set to the `sk-unsloth-...` key, an empty `ANTHROPIC_API_KEY` so Claude Code does not ask for a cloud key, and `ANTHROPIC_MODEL` set to the exact ID returned by `GET /v1/models`. — [source](https://unsloth.ai/docs/basics/claude-code)
- Unsloth's guide says Claude Code "recently prepends" an attribution header (`x-anthropic-billing-header: cc_version=...; cch=...;`) whose value changes on every request, making inference "90% slower with local models", and gives `CLAUDE_CODE_ATTRIBUTION_HEADER=0` (with `CLAUDE_CODE_ENABLE_TELEMETRY=0`) to turn it off. — [source](https://unsloth.ai/docs/basics/claude-code)
- The guide says older Claude Code builds ignored the shell variable `CLAUDE_CODE_ATTRIBUTION_HEADER=0`, so the `--settings` form or the settings-file `env` block is the reliable choice. — [source](https://unsloth.ai/docs/basics/claude-code)
- llama.cpp's Anthropic converter rewrites the five-character `cch=` value of a leading `x-anthropic-billing-header:` system prompt to `fffff`, so the prefix cache survives; the change is llama.cpp PR 21793, commit c807c6e3 dated 2026-04-23. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-chat.cpp)
- The guide's prompt-shrinking flags are `claude --bare --exclude-dynamic-system-prompt-sections`: `--bare` skips auto-discovery of hooks, skills, plugins, MCP servers and CLAUDE.md, and the second flag moves per-machine sections out of the prompt prefix. — [source](https://unsloth.ai/docs/basics/claude-code)
- The guide says to start the server for a coding agent with `unsloth run --model ... --disable-tools --reasoning off -p 8888`, because by default Studio "runs its own server-side tools, which swallows the agent's tool calls, so Claude Code answers but never edits files"; `--disable-tools` switches to passthrough. — [source](https://unsloth.ai/docs/basics/claude-code)
- Server-side tool policy depends on the bind address: tools are on by default on `127.0.0.1` and off by default on `0.0.0.0` or any non-loopback address; `--enable-tools` on `0.0.0.0` shows a y/N prompt (skip with `--yes`) and the resolved policy cannot be overridden per request. — [source](https://unsloth.ai/docs/basics/api)
- Unsloth maps Anthropic `tool_choice` as `auto` to OpenAI `auto`, `any` to `required`, `{type: tool, name}` to `{type: function, function: {name}}`, and `none` to `none`. — [source](https://unsloth.ai/docs/basics/api)
- llama-server's own Anthropic converter maps `tool_choice` `auto` to `auto` and both `any` and `tool` to `required`, dropping the tool name, and does not map `none`. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-chat.cpp)
- Unsloth says server-sent events do not survive the `unsloth studio --secure` Cloudflare quick tunnel and clients must set `stream: false` over it. — [source](https://unsloth.ai/docs/basics/api)
- Unsloth's troubleshooting says `Unable to connect to API (ConnectionRefused)` after returning to the cloud is fixed by `unset ANTHROPIC_BASE_URL`, and a `stream: true` request that returns one JSON blob means the client is buffering or the path is wrong. — [source](https://unsloth.ai/docs/basics/claude-code)
- In April 2026 users reported Studio's inference API as experimental, with `/v1/chat/completions` defaulting to streaming when `stream` is omitted (issue 5047) and tool calling broken through the proxy (issue 4999). — [source](https://github.com/unslothai/unsloth/issues/5141)
- Unsloth PR 5122 extended client-side tool pass-through (from PR 5099) to `/v1/responses` so Codex CLI gets structured `function_call` items instead of raw tool-call tokens in `content`. — [source](https://github.com/unslothai/unsloth/pull/5122)
- Inferred: a Claude Code user on Studio who forgets `--disable-tools` on localhost gets a working chat but no file edits, which looks like a model failure and not a configuration one. — source: `asserted`
