<!-- llms-explorer concept facts · https://llms-explorer.com/tree/claude-code-reactive-compaction-error-wording-re/ · pack 2026-10-05 · ~1248 tokens -->

# Claude Code reactive compaction error-wording recognition

> Claude Code documents three recognised too-long wordings: `Prompt is too long`, Bedrock's `Input is too long for requested model`, and the token `capability_rejected: prompt_too_long` from a Claude apps gateway.

Parent: [Mac local LLMs: Agent clients, context and compaction](https://llms-explorer.com/tree/mac-local-llms-agent-clients-context-and-compaction/) · 1 facets · 17 facts · page: https://llms-explorer.com/tree/claude-code-reactive-compaction-error-wording-re/

## Facts

- Claude Code documents three recognised too-long wordings: `Prompt is too long`, Bedrock's `Input is too long for requested model`, and the token `capability_rejected: prompt_too_long` from a Claude apps gateway. — [source](https://code.claude.com/docs/en/errors)
- Claude Code did not recognise Bedrock's `Input is too long for requested model` before v2.1.217, so auto-compact never triggered on it and `/compact` failed with the same error. — [source](https://code.claude.com/docs/en/errors)
- Claude Code did not recognise `capability_rejected: prompt_too_long` before v2.1.228. — [source](https://code.claude.com/docs/en/errors)
- The gateway troubleshooting table names `ContextWindowExceededError` and `prompt token count of N exceeds the limit of M` as gateway-worded 400s that Claude Code does not recognise, so it does not compact and retry automatically. — [source](https://code.claude.com/docs/en/llm-gateway-connect)
- For a gateway-enforced limit the documented fix is `CLAUDE_CODE_AUTO_COMPACT_WINDOW` set to the gateway limit, which Claude Code clamps to at least 100,000 and at most the model's context window, so a gateway limit below 100,000 cannot be matched and `/compact` stays the recovery. — [source](https://code.claude.com/docs/en/llm-gateway-connect)
- The same table advises setting `CLAUDE_CODE_MAX_OUTPUT_TOKENS` below the gateway model's output limit. — [source](https://code.claude.com/docs/en/llm-gateway-connect)
- llama-server's slot-time overflow message is `request (%d tokens) exceeds the available context size (%d tokens), try increasing it`, raised when the request token count is at least the slot's `n_ctx`, with error type `ERROR_TYPE_EXCEED_CONTEXT_SIZE`. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-context.cpp)
- llama-server formats that error as JSON with `code` 400, the message, and `type` `exceed_context_size_error`. — [source](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/server-common.cpp)
- A user capture shows the same llama-server message in the form `request (4559 tokens) exceeds the available context size (2048 tokens), try increasing it`. — [source](https://github.com/ggml-org/llama.cpp/issues/22746)
- Neither `exceed_context_size_error` nor the llama-server message text appears among the wordings the Claude Code errors page lists as recognised. — [source](https://code.claude.com/docs/en/errors)
- An interactive session shows the too-long condition as `Context limit reached · /compact or /clear to continue`, while `-p` output and the transcript keep `Prompt is too long`; with `DISABLE_COMPACT` set the line names only `/clear`. — [source](https://code.claude.com/docs/en/errors)
- When automatic compaction itself fails, the message is `Prompt is too long · automatic compaction failed: <underlying error>`, and `/compact` fails on the same error until the cause is fixed; before v2.1.229 the cause was omitted. — [source](https://code.claude.com/docs/en/errors)
- Automatic compaction on this error normally summarises the oldest exchanges and keeps the newest; as a last resort it keeps the newest prompt verbatim and summarises the rest, skipped when the carried content has no model reply and under about 1,000 tokens of user text; before v2.1.269 compaction failed whenever it could not summarise a whole exchange. — [source](https://code.claude.com/docs/en/errors)
- A single-exchange conversation is not compacted; Claude Code reports that the request size comes from the system prompt, tool definitions or attachments, giving token counts when the error carries them (before v2.1.162 it attempted compaction anyway). — [source](https://code.claude.com/docs/en/errors)
- Claude Code reports an unrecognised model ID with a `[claude-code:unrecognized_model]` diagnostic line, which a `modelOverrides` entry whose value is the gateway alias silences. — [source](https://code.claude.com/docs/en/model-config)
- Inferred: on a hybrid local model the reactive path is unreliable even with `CLAUDE_CODE_DISABLE_UNKNOWN_MODEL_WINDOW_ENFORCEMENT=1`, so setting `CLAUDE_CODE_AUTO_COMPACT_WINDOW` (at or above the 100,000 clamp) below the server `-c` is the dependable control. — source: `asserted`
- Inferred: a router in front of llama-server could map `exceed_context_size_error` bodies to the text `Prompt is too long` to make the recovery fire, with no documented guarantee. — source: `asserted`
