Claude Code effort levels forwarded to local Anthropic-compatible servers
Parent: Mac local LLMs: Agent clients, context and compaction · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Claude Code resolves the session effort in this order: an explicit choice (`CLAUDE_CODE_EFFORT_LEVEL`, `--effort`, or `/effort`), then saved settings (`modelSettings` per model or the `effortLevel` key), then the model's default effort.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Claude Code resolves the session effort in this order: an explicit choice (`CLAUDE_CODE_EFFORT_LEVEL`, `--effort`, or `/effort`), then saved settings (`modelSettings` per model or the `effortLevel` key), then the model's default effort. [source]
- `CLAUDE_CODE_EFFORT_LEVEL` accepts `low`, `medium`, `high`, `xhigh`, `max` or `auto`, takes precedence over `--effort`, `/effort`, `modelSettings` and `effortLevel`, and a `maxEffortLevel` cap still applies. [source]
- The model-config page lists effort levels only for Claude models (Fable 5.x, Opus 5.x/4.8/4.7, Sonnet 5.x with five levels; Opus 4.6 and Sonnet 4.6 with low, medium, high, max) and says models not listed do not support effort. [source]
- If the level set is unsupported by the active model, Claude Code runs the highest supported level at or below it, for example `xhigh` runs as `high` on Opus 4.6. [source]
- Model default effort is `high` except Opus 5.5 and Sonnet 5.5 (`medium`) and Opus 4.7 (`xhigh`). [source]
- `CLAUDE_CODE_ALWAYS_ENABLE_EFFORT=1` sends the effort parameter on every request even when Claude Code does not recognise the model ID as effort-capable, intended for gateways and providers serving custom identifiers; models that reject effort (Claude 3, Sonnet 4.0 and 4.5, Opus 4.0 and 4.1, Haiku 4.5) stay excluded. [source]
- Claude Code enables effort and extended thinking by matching the model ID against known patterns, so Bedrock ARNs and custom deployment names leave them disabled; `ANTHROPIC_DEFAULT_{OPUS,SONNET,HAIKU,FABLE}_MODEL_SUPPORTED_CAPABILITIES` and `ANTHROPIC_CUSTOM_MODEL_OPTION_SUPPORTED_CAPABILITIES` take a comma list such as `effort,thinking` to declare what a pinned or custom model supports. [source]
- `CLAUDE_CODE_DISABLE_THINKING=1` omits the `thinking` parameter from requests entirely as a compatibility option for proxies that reject it, and on models that think by default the model may still think. [source]
- `MAX_THINKING_TOKENS=0` disables thinking on the Anthropic API but on third-party providers it only omits the `thinking` parameter. [source]
- Claude Code ignores nonzero `MAX_THINKING_TOKENS` on adaptive-reasoning models; the fixed-budget mode (via `CLAUDE_CODE_DISABLE_ADAPTIVE_THINKING=1`) exists only for Opus 4.6 and Sonnet 4.6. [source]
- With thinking turned off on the Anthropic API, Claude Code sends effort `high` instead of a higher level to models that reject that combination, such as Opus 5. [source]
- Claude Code sets `CLAUDE_EFFORT` in Bash tool subprocesses and hook commands, only when the current model supports effort. [source]
- The Anthropic effort page states effort is set with `output_config.effort`, that `adaptive` is a thinking mode and not an effort value, and that changing top-level effort between requests does not preserve cached prefixes. [source]
- Ollama's Anthropic converter sets `think` true for `thinking.type == "enabled"` and false for `"disabled"`; only when `think` is still unset does it read `output_config.effort`. [source]
- Any other `thinking.type` value (for example `adaptive`) leaves Ollama's `think` unset, so effort alone then decides. [source]
- When the model has thinking metadata and the effort string is non-empty, Ollama passes the effort string through as the think level unchanged. [source]
- Without thinking metadata, Ollama trims and lowercases the effort, maps `xhigh` to `high`, and accepts only `low`, `medium`, `high` or `max`; any other value (such as `auto`) is dropped and the model default applies, with no error. [source]
- llama-server's Anthropic-to-OpenAI converter handles `thinking.type == "enabled"` by setting `thinking_budget_tokens` from `budget_tokens` (default 10000) and contains no handling of `output_config` or effort in the fetched file. [source]
- The same converter does not map `thinking.type == "disabled"`, so Claude Code cannot switch thinking off through `/v1/messages` on llama-server; the OpenAI field `reasoning_effort: "none"` is the off switch there. [source]
- llama-server parses `reasoning_effort`: `none` disables reasoning and removes the template kwarg; any other non-empty value is forwarded to the template as `chat_template_kwargs.reasoning_effort`. [source]
- llama.cpp commit 27209a59 (2026-07-24) added `reasoning_effort: "none"` support to the OAI API (#26045). [source]
- llama-server's Responses converter maps only `reasoning.effort` to `reasoning_effort` ("Only effort is handled so far"). [source]
- Unsloth's Claude Code guide launches with `unsloth start claude --model ... --reasoning-effort medium`, and `unsloth run --reasoning on|off` toggles thinking server-side; its settings example also sets `"effortLevel": "high"`. [source]
- Inferred: on a local server whose adapter ignores effort (llama-server), the effective reasoning depth is set only by server flags or template kwargs, so Claude Code's `/effort` slider changes nothing visible. [source]
- Inferred: Claude Code only reaches Ollama's effort mapping for a local model name if it sends effort, which for unrecognised IDs needs `CLAUDE_CODE_ALWAYS_ENABLE_EFFORT=1` or a `_SUPPORTED_CAPABILITIES` declaration. [source]
Children
- No children recorded.