Mac local LLMs: Agent clients, context and compaction
Parent: Running LLM models locally on a Mac · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Needs only ANTHROPIC_BASE_URL + ANTHROPIC_AUTH_TOKEN and a /v1/messages server: Ollama >=0.14 (`ollama launch claude --model <m>`), LM Studio >=0.4.1, oMLX, Rapid-MLX, llama-server; no proxy needed.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Connect Claude Code to a local server
- Needs only ANTHROPIC_BASE_URL + ANTHROPIC_AUTH_TOKEN and a /v1/messages server: Ollama >=0.14 (`ollama launch claude --model <m>`), LM Studio >=0.4.1, oMLX, Rapid-MLX, llama-server; no proxy needed. [source]
- 404 mid-session = background call used the haiku alias: map ANTHROPIC_DEFAULT_HAIKU/SONNET/OPUS_MODEL (and ANTHROPIC_MODEL) to the local id. [source]
- Set CLAUDE_CODE_ATTRIBUTION_HEADER=0 (settings.json env or --settings; old builds ignored shell export) so the billing line stops busting the KV prefix cache. [source]
Subagent model routing (Claude Code)
- Order: per-invocation `model`, then definition frontmatter (`inherit` = main model), then CLAUDE_CODE_SUBAGENT_MODEL, then main model. Before v2.1.251 the variable came first and beat both. Value `inherit` = unset (before v2.1.196 it forced the main model). [source]
- The variable alone does not move built-in Explore/Plan; Explore runs on the `opus` alias model under a gateway via ANTHROPIC_BASE_URL. A user subagent named `Explore` with `model: haiku` overrides it. [source]
- CLAUDE_CODE_SUBAGENT_MODEL_FORCE=1 ignores definition `model` and blocks per-call models; forks and `inherit` skills stay on the main model. blocked values (availableModels) fall back to the inherited model with a warning. [source]
- CLAUDE_CODE_GATEWAY_HINT_HEADERS=1 (v2.1.273+) sends x-claude-code-request-class (main, subagent, workflow, compaction, auxiliary), -agent-type, -prompt-id for router policy. [source]
Windows and compaction (Claude Code)
- Unknown model ID: Claude Code assumes 200k and compacts there. Fix: CLAUDE_CODE_MAX_CONTEXT_TOKENS=<real window>; or CLAUDE_CODE_DISABLE_UNKNOWN_MODEL_WINDOW_ENFORCEMENT=1 (v2.1.223+, compact only after a recognised too-long error). With `[1m]` in the ID also set CLAUDE_CODE_DISABLE_1M_CONTEXT=1. [source]
- CLAUDE_CODE_AUTO_COMPACT_WINDOW: plain integer 100000-1000000; `500k` reads as 500, clamped to 100K. Below 100K use CLAUDE_AUTOCOMPACT_PCT_OVERRIDE (1-100, only lowers). [source]
- Reactive recovery fires only on `Prompt is too long`, `Input is too long for requested model` (>=2.1.217), `capability_rejected: prompt_too_long` (>=2.1.228). Gateway `ContextWindowExceededError` and llama-server `exceed_context_size_error` are not recognised, so no auto-compact. Ollama truncates, never rejects. [source]
Server overflow behaviour
- Ollama (v0.33.1): default num_ctx by VRAM (262144 at >=47 GiB, 32768 at >=23, 4096 below); /v1 cannot set num_ctx, only Modelfile PARAMETER or OLLAMA_CONTEXT_LENGTH. [source]
- llama-server: HTTP 400 `request (N tokens) exceeds the available context size (M tokens), try increasing it`; n_ctx is per slot. `Context size has been exceeded.` is a mid-stream KV failure. --no-context-shift rejects instead of dropping tokens. [source]
Slimming the first request
- Claude Code 2.1.289 first request: default 14.0k; `--tools "Bash,Read,Edit,Write,Grep,Glob"` 3.9k; plus short --system-prompt 2.7k. [source]
- Heavy profile (170-222 MCP tools) sends 66k-96k; `--strict-mcp-config --tools <6>` 10.2k. Env: CLAUDE_CODE_DISABLE_{GIT_INSTRUCTIONS,CLAUDE_MDS,AUTO_MEMORY,NONESSENTIAL_TRAFFIC}=1; CLAUDE_CODE_SIMPLE_SYSTEM_PROMPT=1 keeps tools. [source]
- Non-Anthropic base URL: tool search off, every MCP tool loads; ENABLE_TOOL_SEARCH=true fails if the proxy drops tool_reference blocks. [source]
OpenCode
- First request 6.7k clean, 102.6k with ~/.claude skills (~90k <available_skills>); +89 MCP tools 55.2k. [source]
- Kill skills with per-agent `tools.skill false` or `permission.skill {"*":"deny"}`; global `tools.skill false` and experimental.mcp_lazy did nothing in 1.18.34. [source]
- `"tools": {"*": false}` plus per-agent enables gives 9 tools, 3,260 tokens. Compaction V2: keep.tokens 15000, buffer 10%; reserved/prune are V1. Set limit.context to real n_ctx. [source]
Codex
- Unknown slug warns `Model metadata for <slug> not found`; set model_context_window and model_catalog_json. Auto-compact at 90% of window. [source]
- Symptom: session shows no `mcp__*` tools and no `tool_search` though `codex mcp list` is fine. Cause: catalog entry `supports_search_tool: true` (DeepSeek's setup script hardcodes it with `tool_mode: null`) defers every MCP tool, and `tool_mode: null` has no code-mode exec to hold them. Fix: set it false in ~/.codex/models.json (3 reports, 0.145-0.147). [source]
- Correction: earlier claim that deferral needs `supports_search_tool` AND provider `namespace_tools` is wrong on main (2026-10-05): spec_plan.rs gates on `supports_search_tool` alone, and the string `namespace_tools` appears in no provider file; `ToolSpec::Namespace` serializes with no provider condition. [source]
- Other exposure rules: false makes MCP tools direct unless listed in per-server `omit_tools_from`; a registered tool named `tool_search` is removed; v1 `spawn_agent` and multi_agent_version "v1" tools go Deferred when true, so a model that never calls `tool_search` cannot reach them. [source]
- Via LiteLLM, `type: "tool_search"` gives `400: Field required ... tools[13].function`; rewriting it gives `tool_search handler received unsupported payload`; false still hid MCP tools. Dropped `reasoning_content` gives `400: The reasoning_content in the thinking mode must be passed back to the API`. [source]
- Provider capabilities (codex-model-provider): `remote_compaction` defaults `Unsupported` (local compaction); `V2` (sends `compaction_trigger` items) is set only for OpenAI/Azure Responses providers. Config may override only `external_web_access` and `remote_compaction` (unknown fields rejected); [source]
- Responses `tool_search` is gpt-5.4+ only. Hosted: `execution: "server"`, `call_id: null`. Client: tool with `execution: "client"`; answer with `tool_search_output` (`execution: "client"`, `call_id`, `status: completed`, `tools`). A local model needs a proxy emulating the items. [source]
- MCP as `namespace` tools: llama.cpp skips silently, LM Studio rejects, flat names fail `unsupported call`; codex-universal-proxy (formerly codex-ollama-proxy) flattens. [source]
Open questions
- Unmeasured: success vs tool count on 9B-35B models; which llama-server error text Claude Code recovers from; which Codex commit removed `namespace_tools` and whether a Chat-Completions bridge gate exists. [source]
Corrections and disagreements
- Ollama default context: Aider docs still say Ollama defaults to 2k; Ollama FAQ/ existing playbook say 4k/32k/256k by VRAM; Kulman saw 32768 on 48 GB. CONTRADICTS: aider.chat/docs/llms/ollama.html vs existing local-llm-troubleshooting-playbook.md section 2.1 (the playbook is right for current Ollama). [source]
- CONTRADICTS: claude-code-system-prompt-and-tool-schema-slimming-for-local-models.md says `OPENCODE_DISABLE_CLAUDE_CODE=1` alone did not remove the 90k listing and that disabling the `skill` tool did. Both are true here but incomplete: skills were also read from `~/.agents/skills` and `~/.config/opencode/skills`, which that flag does not cover (the first 3 listed paths in the request were `~/.agents/skills/...`), and the disable that works is per-agent `tools.skill false` or `permission.skill "*":"deny"`; the global `tools` key is ignored for this purpose in 1.18.34. [source]
- CONTRADICTS: claude-code-context-window-mismatch-against-local-servers.md lists reserved and prune as the compaction config; that is the V1 schema, and a V2 install uses keep.tokens and buffer. [source]
- This CONTRADICTS the inferred claim in codex-responses-api-compatibility-on-local-serve.md that it is unknown which functions a skipped `namespace` carries: Codex wraps each configured MCP server as `{"type":"namespace","name":"mcp__<server>__","tools":[function, ...]}`, so skipping it removes every tool of that MCP server. [source]
- CONTRADICTS: codex-namespace-and-hosted-tool-types-skipped-by.md on the proxy's name: `codex-ollama-proxy` is now `codex-universal-proxy`; the first universal command migrates `~/.codex/ollama-shape-proxy` to `~/.codex/codex-universal-proxy`, copies legacy catalog and reference files forward, and replaces legacy service registrations. [source]
- This CONTRADICTS the unqualified wording in local-router-policy-keyed-on-x-claude-code-reque.md that the variable sets the default for subagents with a short override note: the full order is per-invocation `model`, then definition `model` frontmatter (`inherit` selects the main model), then `CLAUDE_CODE_SUBAGENT_MODEL`, then the main conversation's model. [source]
- CONTRADICTS: codex-tool-search-deferred-mcp-tools-and-the-sup.md (line 10, `search_tool_enabled` computed in spec_plan.rs as `supports_search_tool && provider.capabilities().namespace_tools`): on main fetched 2026-10-05 `spec_plan.rs` contains no `namespace_tools` and gates MCP deferral on `model_info.supports_search_tool` alone. [source]
Concepts in this cluster
- Local LLM server as a coding-agent backend on Mac [source]
- Claude Code context-window mismatch against local servers [source]
- Claude Code system-prompt and tool-schema slimming for local models [source]
- claude-code-router Anthropic-to-OpenAI translation over a local gateway [source]
- Agent skill-listing injection cost [source]
- Claude Code unknown-model context window handling [source]
- Task-success versus tool count and prompt size for small local models [source]
- Claude Code helper-model request to custom base URL [source]
- Claude Code skill listing budget formula [source]
- Hook additional-context injection into the first message [source]
- MCP tool-definition token cost in OpenCode [source]
- Multi-root skill discovery and duplicated skill directories [source]
- Claude Code compaction floor of 100K versus small local windows [source]
- OpenCode compaction config for local providers [source]
- mcp-compressor wrapper proxy in front of a local-model agent [source]
- Claude Code effort levels forwarded to local Anthropic-compatible servers [source]
- Client-side compaction thresholds for hybrid-model local servers [source]
- Claude Code gateway hint headers for local routers [source]
- Claude Code reactive compaction error-wording recognition [source]
- Codex auto-compaction body_after_prefix scope [source]
- Codex namespace and hosted tool types skipped by local servers [source]
- llama.cpp normalisation of Claude Code billing header cch [source]
- Codex model catalog JSON (model_catalog_json) for local model slugs [source]
- Flattening proxies for Codex namespace tools (codex-ollama-proxy and similar) [source]
- Local router policy keyed on x-claude-code-request-class [source]
- Rewriting llama-server overflow errors to Claude Code's recognised wording [source]
- Codex tool_search deferred MCP tools and the supports_search_tool catalog flag [source]
- Responses API client-executed tool_search support in local servers [source]
- Router session keying when subagents reuse metadata.user_id [source]
- CLAUDE_CODE_SUBAGENT_MODEL routing of subagents to a small local model [source]
- Codex model-visible tool exposure rules (spec_plan.rs namespace_tools capability [source]
- Codex provider capabilities struct and where namespace serialization is gated [source]
Children
- mcp-compressor wrapper proxy in front of a local-model agent
- MCP tool-definition token cost in OpenCode
- Multi-root skill discovery and duplicated skill directories
- OpenCode compaction config for local providers
- Responses API client-executed tool_search support in local servers
- Rewriting llama-server overflow errors to Claude Code's recognised wording
- Router session keying when subagents reuse metadata.user_id
- Task-success versus tool count and prompt size for small local models
- Agent skill-listing injection cost
- Claude Code compaction floor of 100K versus small local windows
- Claude Code context-window mismatch against local servers
- Claude Code effort levels forwarded to local Anthropic-compatible servers
- Claude Code gateway hint headers for local routers
- Claude Code helper-model request to custom base URL
- Claude Code reactive compaction error-wording recognition
- claude-code-router Anthropic-to-OpenAI translation over a local gateway
- Claude Code skill listing budget formula
- CLAUDE_CODE_SUBAGENT_MODEL routing of subagents to a small local model
- Claude Code system-prompt and tool-schema slimming for local models
- Claude Code unknown-model context window handling
- Client-side compaction thresholds for hybrid-model local servers
- Codex auto-compaction body_after_prefix scope
- Codex model catalog JSON (model_catalog_json) for local model slugs
- Codex model-visible tool exposure rules (spec_plan.rs namespace_tools capability
- Codex namespace and hosted tool types skipped by local servers
- Codex provider capabilities struct and where namespace serialization is gated
- Codex tool_search deferred MCP tools and the supports_search_tool catalog flag
- Flattening proxies for Codex namespace tools (codex-ollama-proxy and similar)
- Hook additional-context injection into the first message
- llama.cpp normalisation of Claude Code billing header cch
- Local LLM server as a coding-agent backend on Mac
- Local router policy keyed on x-claude-code-request-class