Client-side compaction thresholds for hybrid-model local servers
Parent: Mac local LLMs: Agent clients, context and compaction · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
With no window set, Claude Code compacts when the conversation reaches the model's context limit, except cloud sessions, Sonnet 4.6 and Opus 4.6 without extended context (200K boundary), Opus 4.8 and later on a 200K window, and `CLAUDE_CODE_DISABLE_1M_CONTEXT=1`.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- With no window set, Claude Code compacts when the conversation reaches the model's context limit, except cloud sessions, Sonnet 4.6 and Opus 4.6 without extended context (200K boundary), Opus 4.8 and later on a 200K window, and `CLAUDE_CODE_DISABLE_1M_CONTEXT=1`. [source]
- Models with a native 1M window compact before the window fills, at about 967K tokens by default. [source]
- `/autocompact` and `--autocompact` accept 100K to 1M as a plain count, a `k` or `M` suffix, or a bare number from 100 to 1000 meaning thousands; `CLAUDE_CODE_AUTO_COMPACT_WINDOW` accepts only a plain integer. [source]
- Precedence is `CLAUDE_CODE_AUTO_COMPACT_WINDOW` over `/autocompact` over the `--autocompact` flag over `autoCompactWindow`; the flag is not preempted by managed settings but the command is. [source]
- A window saved with `/autocompact` is stored per model under `modelSettings` from v2.1.288; earlier versions saved one window for every model. [source]
- Once `CLAUDE_CODE_AUTO_COMPACT_WINDOW` is set, the status line `used_percentage` still measures against the model's full window and no longer indicates when compaction will run. [source]
- `CLAUDE_AUTOCOMPACT_PCT_OVERRIDE` can only lower the threshold (values above the default percentage are ignored), applies only in sessions that compact before the model's context limit, and covers both main conversations and subagents. [source]
- On a model ID Claude Code does not recognise, `CLAUDE_CODE_DISABLE_UNKNOWN_MODEL_WINDOW_ENFORCEMENT=1` makes it compact only after the API rejects the conversation with a too-long error it recognises, and Claude Code does not run that recovery when a gateway rewrites the error wording. [source]
- With `CLAUDE_CODE_GATEWAY_HINT_HEADERS=1` (default off on a custom base URL, v2.1.273 or later) Claude Code sends `x-claude-code-request-class` whose value `compaction` marks the summarisation request. [source]
- `x-claude-code-compaction` appears only on the summarising request and carries the trigger: `auto` (window near capacity), `manual` (`/compact`) or `reactive` (the API rejected a request as too long). [source]
- `x-claude-code-context-compacted` appears once, on the first main-conversation request after a compaction, and signals that the prefix before this request is no longer used so a cache keyed on it can be dropped. [source]
- Setting `CLAUDE_CODE_GATEWAY_HINT_HEADERS=0` stops the headers on every connection, including the default send to the Anthropic API. [source]
- Codex's `model_auto_compact_token_limit` is the token threshold that triggers automatic history compaction (unset uses model defaults), and `model_context_window` sets the window Codex assumes for the active model. [source]
- Codex's `model_auto_compact_token_limit_scope` is `total` (default; counts the full active context) or `body_after_prefix` (counts only growth after the carried compaction-window prefix). [source]
- A Codex CLI request through a llama.cpp backend was 18383 tokens at the first turn and failed against an 8192-token server window with `exceed_context_size_error`; raising `--ctx-size` to 32768 fixed it. [source]
- A user report puts Codex CLI's system prompt at "at least 27,000 tokens" and runs llama-server with `-c 32768`; Ollama's guidance for Codex is a 64k context and LM Studio's is above about 25k. [source]
- Ollama's guidance for Codex is "at least 64k tokens" and LM Studio's is "more than ~25k" per a September 2026 how-to. [source]
- Inferred: a local router that sees `x-claude-code-context-compacted` can evict the conversation's recurrent-state checkpoints at once instead of waiting for LRU, since the old prefix will not recur. [source]
- Inferred: because a hybrid server cannot shift context, the compaction window should be set at or below the server's `-c` minus the largest expected single tool result, not at the model's trained maximum. [source]
Children
- No children recorded.