Claude Code reactive compaction error-wording recognition
Parent: Mac local LLMs: Agent clients, context and compaction · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Claude Code documents three recognised too-long wordings: `Prompt is too long`, Bedrock's `Input is too long for requested model`, and the token `capability_rejected: prompt_too_long` from a Claude apps gateway.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Claude Code documents three recognised too-long wordings: `Prompt is too long`, Bedrock's `Input is too long for requested model`, and the token `capability_rejected: prompt_too_long` from a Claude apps gateway. [source]
- Claude Code did not recognise Bedrock's `Input is too long for requested model` before v2.1.217, so auto-compact never triggered on it and `/compact` failed with the same error. [source]
- Claude Code did not recognise `capability_rejected: prompt_too_long` before v2.1.228. [source]
- The gateway troubleshooting table names `ContextWindowExceededError` and `prompt token count of N exceeds the limit of M` as gateway-worded 400s that Claude Code does not recognise, so it does not compact and retry automatically. [source]
- For a gateway-enforced limit the documented fix is `CLAUDE_CODE_AUTO_COMPACT_WINDOW` set to the gateway limit, which Claude Code clamps to at least 100,000 and at most the model's context window, so a gateway limit below 100,000 cannot be matched and `/compact` stays the recovery. [source]
- The same table advises setting `CLAUDE_CODE_MAX_OUTPUT_TOKENS` below the gateway model's output limit. [source]
- llama-server's slot-time overflow message is `request (%d tokens) exceeds the available context size (%d tokens), try increasing it`, raised when the request token count is at least the slot's `n_ctx`, with error type `ERROR_TYPE_EXCEED_CONTEXT_SIZE`. [source]
- llama-server formats that error as JSON with `code` 400, the message, and `type` `exceed_context_size_error`. [source]
- A user capture shows the same llama-server message in the form `request (4559 tokens) exceeds the available context size (2048 tokens), try increasing it`. [source]
- Neither `exceed_context_size_error` nor the llama-server message text appears among the wordings the Claude Code errors page lists as recognised. [source]
- An interactive session shows the too-long condition as `Context limit reached · /compact or /clear to continue`, while `-p` output and the transcript keep `Prompt is too long`; with `DISABLE_COMPACT` set the line names only `/clear`. [source]
- When automatic compaction itself fails, the message is `Prompt is too long · automatic compaction failed: <underlying error>`, and `/compact` fails on the same error until the cause is fixed; before v2.1.229 the cause was omitted. [source]
- Automatic compaction on this error normally summarises the oldest exchanges and keeps the newest; as a last resort it keeps the newest prompt verbatim and summarises the rest, skipped when the carried content has no model reply and under about 1,000 tokens of user text; before v2.1.269 compaction failed whenever it could not summarise a whole exchange. [source]
- A single-exchange conversation is not compacted; Claude Code reports that the request size comes from the system prompt, tool definitions or attachments, giving token counts when the error carries them (before v2.1.162 it attempted compaction anyway). [source]
- Claude Code reports an unrecognised model ID with a `[claude-code:unrecognized_model]` diagnostic line, which a `modelOverrides` entry whose value is the gateway alias silences. [source]
- Inferred: on a hybrid local model the reactive path is unreliable even with `CLAUDE_CODE_DISABLE_UNKNOWN_MODEL_WINDOW_ENFORCEMENT=1`, so setting `CLAUDE_CODE_AUTO_COMPACT_WINDOW` (at or above the 100,000 clamp) below the server `-c` is the dependable control. [source]
- Inferred: a router in front of llama-server could map `exceed_context_size_error` bodies to the text `Prompt is too long` to make the recovery fire, with no documented guarantee. [source]
Children
- No children recorded.