Prompt caching
Parent: LLM Models and APIs · Published reference · snapshot 2026-08-31
↓ Facts as markdownall context files
Prompt caching lets the API reuse an already-processed prompt prefix (tools → system → messages) across requests to cut cost and latency: Anthropic offers automatic caching or explicit `cache_control: {"type": "ephemeral"}` breakpoints with 5-minute or 1-hour TTLs, bills cache writes at 1.25×/2× and cache hits at 0.1× the base input price, and reports them in usage.cache_creation_input_tokens / cache_read_input_tokens; OpenRouter passes cache_control through to Anthropic and Gemini, adds provider sticky routing, and documents OpenAI automatic caching. Two of the three docsets contributed (the OpenAI export is empty).
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Definitions
- Automatic caching — Automatic caching is the simplest way to enable prompt caching. Instead of placing `cache_control` on individual content blocks, add a single `cache_control` field at the top level of your request body. [source]
- Prompt caching — Prompt caching doesn't reduce the number of tokens in context, but it reduces what you pay for them on subsequent requests. If your tool definitions are stable, cache them once and reuse the cached prefix across thousands of requests. [source]
- Cache miss reason types — `cache_miss_reason` is a discriminated union on `type`. The response reports the earliest divergence only, so fix it first; later ones may be hidden behind it. [source]
- Caching — Some endpoints/models provide implicit caching of prompts. This keeps repeated prompt data in an in-memory cache in the provider's datacenter, so that the repeated part of the prompt does not need to be re-processed. [source] — OpenRouter's description of provider implicit caching
Structure and components
- Cache storage and sharing — Prompt caching uses [workspace](https://platform.claude.com/docs/en/manage-claude/workspaces)-level isolation. Caches are isolated per workspace, ensuring data separation between workspaces within the same organization. [source]
- Advisor prompt caching — There are two independent caching layers. [source]
How it works
- How it works — Set `max_tokens: 0` in your request. The API reads your prompt into the model and writes the cache at any `cache_control` breakpoint, then returns immediately without generating any output. [source]
- How automatic caching works in multi-turn conversations — With automatic caching, the cache point moves forward automatically as conversations grow. Each new request caches everything up to the last cacheable block, and previous content is read from cache. [source]
- Explicit cache breakpoints — For more control over caching, you can place `cache_control` directly on individual content blocks. This is useful when you need to cache different sections that change at different frequencies, or need fine-grained control over exactly what gets cached. [source]
- Server tool results are cached automatically — When your request has prompt caching enabled and Claude uses a [server tool](https://platform.claude.com/docs/en/agents-and-tools/tool-use/server-tools) such as web search, web fetch, or code execution, the API automatically places a cache breakpoint on the server tool result before running the… [source]
- Implicit Caching — Gemini 2.5 series models and newer support **implicit caching**, providing automatic caching functionality similar to OpenAI’s automatic caching. Implicit caching works seamlessly — no manual setup or additional `cache_control` breakpoints required. [source]
- Provider Sticky Routing — To maximize cache hit rates, OpenRouter uses **provider sticky routing** to route your subsequent requests to the same provider endpoint after a cached request. This works automatically with both implicit caching (e.g. [source]
- Combining with block-level caching — Automatic caching is compatible with [explicit cache breakpoints](https://platform.claude.com/docs/en/build-with-claude/prompt-caching#explicit-cache-breakpoints). When used together, the automatic cache breakpoint uses one of the 4 available breakpoint slots. [source]
- When to use a mid-conversation system message — [Prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching) hashes the request prefix in order: `tools`, then `system`, then `messages`. A cache hit requires the prefix to match a recent request exactly, byte for byte, up to the cache breakpoint. [source]
- Caching in the Batch API — `cache_control` breakpoints work on Anthropic `:batch` endpoints the same way as on the sync API, but the requests inside a single batch may process concurrently and in any order — a cache written by one line is not guaranteed to be visible to other lines in the same batch. [source]
- Prompt caching — Compaction works well with [prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching). You can add a `cache_control` breakpoint on compaction blocks to cache the summarized content. [source]
- Pre-warming the cache — Cache pre-warming lets you load your system prompt or tool definitions into the prompt cache before a user triggers a real request. This eliminates the cache-miss latency penalty on the first user interaction, reducing time-to-first-token (TTFT) for latency-sensitive applications. [source]
- Executor-side caching — The `advisor_tool_result` block is cacheable like any other content block. A `cache_control` breakpoint placed after it on a subsequent turn hits. [source]
- What invalidates your cache — The cache follows a prefix hierarchy (`tools` → `system` → `messages`), so a change at one level invalidates that level and everything after it: [source]
- Prompt Caching — Discovering a tool does not disturb the tools already in the conversation, so prompt caching is preserved across a search. You can start a conversation with a small loaded set, let the model discover more as it goes, and keep your cache hit across every turn. [source]
- Using prompt caching with Message Batches — The Message Batches API supports prompt caching, allowing you to potentially reduce costs and processing time for batch requests. The pricing discounts from prompt caching and Message Batches can stack, providing even greater cost savings when both features are used together. [source]
- Prompt caching in manual mode — Manual mode adds one rule on top of the mode-neutral caching behavior described in [thinking and prompt caching](https://platform.claude.com/docs/en/build-with-claude/thinking#thinking-and-prompt-caching): changing `budget_tokens` between requests invalidates cache breakpoints, just as switching… [source]
- [Tool search](https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-search-tool): Discovered tools load as `tool_reference` blocks, preserving prefix cache [source]
- defer\_loading and cache preservation — Deferred tools are not included in the system-prompt prefix. When the model discovers a deferred tool through [tool search](https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-search-tool), the definition is appended inline as a `tool_reference` block in the conversation history. [source]
- Cache-aware ITPM — Many API providers use a combined "tokens per minute" (TPM) limit that may include all tokens, both cached and uncached, input and output. [source]
- Prompt caching — Consecutive requests that keep the same thinking configuration and effort level preserve prompt caching; see [Thinking and prompt caching](https://platform.claude.com/docs/en/build-with-claude/thinking#thinking-and-prompt-caching) for the full rules. [source]
- `defer_loading` and prompt caching — Tools with `defer_loading: true` are stripped from the rendered tools section before the cache key is computed. They don't appear in the system-prompt prefix at all. [source]
- Mid-conversation tool changes (beta) — You can add or remove tools between turns of a conversation while preserving the prompt cache, instead of resending a fixed tool list for the life of a session. Mid-conversation tool changes are in beta: include the `mid-conversation-tool-changes-2026-07-01` beta header in your requests. [source]
- Session Stickiness — The Pareto Router reuses the selected **model** and **provider** on a best-effort basis so that subsequent requests in the same conversation route to the same place. This keeps behavior consistent within a conversation and maximizes [prompt cache](/docs/guides/best-practices/prompt-caching) hits. [source] — OpenRouter Pareto Router
- How cache diagnostics works — When the beta header is present, the API stores a lightweight fingerprint of each request, keyed by the response `id`. On your next request, include that `id` as `diagnostics.previous_message_id`. [source]
Parameters and configuration
- [Automatic prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching#automatic-caching): Description=Simplify prompt caching to a single API parameter. The system automatically caches the last cacheable block in your request, moving the cache point forward as conversations grow.; ZDR=ZDR eligible [source]
- `null`: Cache read tokens=low or zero; Interpretation=Your requests match but the cache entry was no longer available. Consider shortening gaps between turns or using the [1-hour cache TTL](https://platform.claude.com/docs/en/build-with-claude/prompt-caching#1-hour-cache-duration). [source]
- `cache_miss_reason` is a `*_changed` type: Cache read tokens=high; Interpretation=Rare. A change occurred late in the prompt but an earlier `cache_control` breakpoint still hit. Worth fixing, but low impact. [source]
- Every request is a cache miss: Likely cause=`tool_choice`, the thinking configuration, or `output_config.effort` varying between requests; Fix=Keep `tool_choice` stable or place the `cache_control` breakpoint before the variation point; hold the thinking configuration and effort level constant for the life of a cached conversation. See [Tool use with prompt… [source]
- `cache_control`: Required=No; Description=[Prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching) breakpoint at the toolset definition; entry only. A breakpoint on any `tool_use` or `tool_result` block in a batch takes effect at the end of that batch; see [Tool use with prompt… [source]
- `cache_control`: Type=object; Required=No; Description=[Prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching) cache breakpoint configuration for this toolset. [source]
- `cache_control`: Purpose=Set a prompt-cache breakpoint at this tool definition; Available on=All tools (on `computer_toolset_20260801` and `browser_toolset_20260801`, set it on the toolset entry itself, not inside member `configs`); Detailed guide=[Prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching) [source]
- [Prompt caching (1hr)](https://platform.claude.com/docs/en/build-with-claude/prompt-caching#1-hour-cache-duration): Description=Extended 1-hour cache duration for less frequently accessed but important context, complementing the standard 5-minute cache.; ZDR=ZDR eligible [source]
- [Prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching): Endpoint=`/v1/messages`; ZDR eligible=<Eligible>Yes</Eligible>; HIPAA eligible=<Eligible>Yes</Eligible>; Details=Your prompts and Claude's outputs are not stored. KV cache representations and cryptographic hashes are held in memory for the cache TTL and promptly deleted after expiry. See [Prompt… [source]
- `cache_miss_reason` is a `*_changed` type: Cache read tokens=low or zero; Interpretation=Your bug. The request changed; fix the cause indicated by `type`. [source]
- 1-hour cache duration — If you find that 5 minutes is too short, Anthropic also offers a 1-hour cache duration [at additional cost](https://platform.claude.com/docs/en/build-with-claude/prompt-caching#pricing). [source]
- `null`: Cache read tokens=high; Interpretation=Working as expected. Your prefix is stable and the cache hit. [source]
- Usage object fields — The usage object in API responses includes detailed cache metrics in the `prompt_tokens_details` field: [source]
- [Prompt caching (5m)](https://platform.claude.com/docs/en/build-with-claude/prompt-caching): Description=Provide Claude with more background knowledge and example outputs to reduce costs and latency.; ZDR=ZDR eligible [source]
- `system_changed`: What it means=The `system` parameter differs. Typically a timestamp, request ID, or other per-request value was interpolated into the system prompt.; What to change=Make the system prompt a byte-stable constant and move dynamic data into the first `user` message after your cache breakpoint. [source]
- `cache_control`: Type=object; Description=Cache control settings (for example, `{"type": "ephemeral"}`) [source]
- Tracking cache performance — Monitor cache performance using these API response fields, within `usage` in the response (or `message_start` event if [streaming](https://platform.claude.com/docs/en/build-with-claude/streaming)): [source]
- `caching`: Type=object \| null; Default=`null` (off); Description=Enables [prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching) for the advisor's own transcript across calls within a conversation. See [Advisor prompt caching](https://platform.claude.com/docs/en/agents-and-tools/tool-use/advisor-tool#advisor-prompt-caching). [source]
- Configure the toolset — Besides `type`, the toolset entry accepts `configs`, `cache_control`, and `allowed_callers`; the rules these fields share with the computer use toolset are listed under [Client toolsets](https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-reference#client-toolsets), and this section… [source]
- The toolset entry accepts configs, cache_control, and allowed_callers fields in addition to type. [source]
- `null`: Either `previous_message_id` was `null` (first turn, nothing to compare), or a comparison ran and found no divergence. [source]
- `{"cache_miss_reason": null}`: The comparison was still running when the response was serialized. This can happen when the response starts very quickly. Treat it as inconclusive and check the next turn. [source]
- `{"cache_miss_reason": {...}}`: A `cache_miss_reason` is attached. For `*_changed` types this identifies the first divergence point; `previous_message_not_found` and `unavailable` are cases where no comparison was produced. [source]
- `model_changed`: What it means=The `model` differs from the previous request (for example, a router, A/B test, or fallback selected a different model). The cache is per-model.; What to change=Hold the model constant within a cached conversation. [source]
- `tools_changed`: What it means=The `tools` array differs: tools were added, removed, or reordered between turns, or tool `input_schema` JSON was serialized non-deterministically.; What to change=Send the same tool list on every turn in a fixed order with deterministically serialized schemas (for example, sort keys). [source]
- `messages_changed`: What it means=The model, system, and tools all match, but an earlier entry in `messages` was altered, reordered, or removed rather than appended to. Typically conversation history was truncated or edited, or assistant turns and `tool_result` blocks were re-serialized differently on resend.; What to change=Treat the history as append-only; echo assistant `content` and tool… [source]
- `previous_message_not_found`: What it means=No stored fingerprint exists for the supplied `previous_message_id`. This is not evidence that your request changed. Typically the previous request did not carry the beta header, it came from a different workspace, or too much time has passed since it was sent.; What to change=Send the beta header on every turn and keep consecutive turns close together… [source]
- Tool definition properties — Every tool in the `tools` array, including user-defined tools, accepts optional properties that control how the tool is loaded, who can call it, and how its inputs are validated. These properties compose: you can set `defer_loading` and `cache_control` and `strict` on the same tool. [source]
- [Cache diagnostics](https://platform.claude.com/docs/en/build-with-claude/cache-diagnostics): Endpoint=`/v1/messages` (with `diagnostics`); ZDR eligible=<Eligible status="qualified">Yes (qualified)</Eligible>; HIPAA eligible=<Eligible status="no">No</Eligible>; Details=Your prompts and Claude's outputs are not stored. A fingerprint of cryptographic hashes and token-count estimates is retained… [source]
How-to and procedures
- Mixing different TTLs — You can use both 1-hour and 5-minute cache controls in the same request, but with an important constraint: Cache entries with longer TTL must appear before shorter TTLs (that is, a 1-hour cache entry must appear before any 5-minute cache entries). [source]
- cache\_control on tool definitions — Place `cache_control: {"type": "ephemeral"}` on the last tool in your `tools` array. This caches the entire tool-definitions prefix, from the first tool through the marked breakpoint: [source]
- Consider using the 1-hour cache duration with prompt caching for batches with shared context to improve cache hit rates, as processing can take longer than 5 minutes. [source]
- Structuring your prompt — Place static content (tool definitions, system instructions, context, examples) at the beginning of your prompt. Mark the end of the reusable content for caching using the `cache_control` parameter. [source]
- How to Enable Gemini Prompt Caching: — Gemini caching in OpenRouter requires you to insert `cache_control` breakpoints explicitly within message content, similar to Anthropic. [source]
- Use prompt caching to reduce cost and latency by caching prompt prefixes shared across requests in a batch. [source]
- When to use the 1-hour cache — If you have prompts that are used at a regular cadence (that is, system prompts that are used more frequently than every 5 minutes), continue to use the 5-minute cache, because this will continue to be refreshed at no additional charge. [source]
- Using `session_id` for sticky sessions — For more explicit control over sticky routing, you can pass a `session_id` in your request. When a `session_id` is present, OpenRouter uses it directly as the sticky routing key instead of deriving one from message hashing. [source]
- Modifying tool results — You can modify tool results before they're sent back to Claude. This is useful for adding metadata such as `cache_control` to enable [prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching) on tool results, or for transforming the tool output. [source]
- **Shrink Tool-Heavy Prefixes**: Recommended Configuration=`openrouter:tool_search` + `defer_loading: true`; Latency Impact=Defer most tool definitions out of the prompt; keeps cacheable prefix lean and stable across turns. [source]
- Reading diagnostics alongside usage — `diagnostics` answers "did my request change?" while `usage.cache_read_input_tokens` answers "did the cache hit?". Combining them tells you where to look. [source]
- Combining approaches — These approaches compose. A long-running agent might use tool search to keep the toolset lean, prompt caching to amortize the cost of the remaining definitions, and context editing to trim stale results as the conversation grows. [source]
- Cut spend without losing quality — Prompt caching, token hygiene, batch processing, and a prompt audit against your current model all lower what you pay without lowering output quality. [source]
- **Agent Loops / Long Contexts**: Recommended Configuration=Stable `session_id` + static prompt prefix; Latency Impact=Slashes prefill time and cost (up to 80–90%) via provider KV cache. [source]
- System messages must be placed after the user turn to keep cached bytes untouched and satisfy placement rules. [source]
- Advisor-side caching — Set `caching` on the tool definition to enable prompt caching for the advisor's own transcript across calls within the same conversation: [source]
- Any workload, any model: Do this=Turn on prompt caching and trim unneeded tokens; both are free; Where=[Cache repeated context](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#cache-repeated-context) · [Trim tokens](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#trim-input-and-context-tokens) [source]
- Basic usage — Send the beta header on every turn. On the first turn, pass `"previous_message_id": null` to opt in without a prior message to compare against. [source]
- Threading diagnostics through a conversation loop — In a multi-turn conversation, carry the latest response `id` forward as `previous_message_id` on every turn. The first iteration passes `null` to opt in; each subsequent iteration passes the `id` from the previous response. [source]
- Cache control — Add `cache_control` on the search result block to cache it for reuse across requests. It sits alongside `citations` on the same block: [source]
- A person waits between turns: Do this=Use the 1-hour cache duration; it is cheaper once about 1 turn in 20 follows a pause between 5 minutes and an hour and few gaps run over an hour; Where=[Pick the cache duration](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#pick-the-cache-duration) [source]
- User tracking — Track your end-users by including a `user` parameter in API requests. This improves caching performance (sticky routing per user) and enables user-level analytics in your activity feed and exports. [source]
Examples and snippets
- TTL support: { "cache_control": { "type": "ephemeral", "ttl": "1h" } } — `{ "cache_control": { "type": "ephemeral", "ttl": "1h" } }` [source]
- 1-hour cache duration: "cache_control": { — `"cache_control": {` [source]
- cURL: curl https://api.anthropic.com/v1/messages \ — `# Warm the cache at application startup or on a scheduled interval.` [source] — pre-warm request; CLI variant at the same anchor
- Effort changes invalidate the prompt cache: import requests — `import requests` [source] — example; 7 language variants at the same anchor
- Request 1: Content=System + User(1) + Asst(1) + **User(2)** ◀ cache; Cache behavior=Everything written to cache [source]
- Request 2: Content=System + User(1) + Asst(1) + User(2) + Asst(2) + **User(3)** ◀ cache; Cache behavior=System through User(2) read from cache; Asst(2) + User(3) written to cache [source]
- Request 3: Content=System + User(1) + Asst(1) + User(2) + Asst(2) + User(3) + Asst(3) + **User(4)** ◀ cache; Cache behavior=System through User(3) read from cache; Asst(3) + User(4) written to cache [source]
- Prompt caching examples — To help you get started with prompt caching, the [prompt caching cookbook](https://platform.claude.com/cookbook/misc-prompt-caching) provides detailed examples and best practices. [source]
- lines: Is prompt caching saving me anything? Show prompt vs completion — `Is prompt caching saving me anything? Show prompt vs completion` [source] — OpenRouter analytics recipe
Measurements and reference values
- Lower prompt cache minimum — The minimum cacheable prompt length on Claude Opus 5 is 512 tokens, down from 1,024 tokens on Claude Opus 4.8. Prompts that were too short to cache on Claude Opus 4.8 can now create cache entries with no code changes. [source]
- Claude Fable 5: Base Input Tokens=$10 / MTok; 5m Cache Writes=$12.50 / MTok; 1h Cache Writes=$20 / MTok; Cache Hits & Refreshes=$1 / MTok; Output Tokens=$50 / MTok [source]
- Claude Mythos 5 ([limited availability](https://anthropic.com/glasswing)): Base Input Tokens=$10 / MTok; 5m Cache Writes=$12.50 / MTok; 1h Cache Writes=$20 / MTok; Cache Hits & Refreshes=$1 / MTok; Output Tokens=$50 / MTok [source]
- Claude Opus 5: Base Input Tokens=$5 / MTok; 5m Cache Writes=$6.25 / MTok; 1h Cache Writes=$10 / MTok; Cache Hits & Refreshes=$0.50 / MTok; Output Tokens=$25 / MTok [source]
- Claude Opus 4.8: Base Input Tokens=$5 / MTok; 5m Cache Writes=$6.25 / MTok; 1h Cache Writes=$10 / MTok; Cache Hits & Refreshes=$0.50 / MTok; Output Tokens=$25 / MTok [source]
- Claude Opus 4.7: Base Input Tokens=$5 / MTok; 5m Cache Writes=$6.25 / MTok; 1h Cache Writes=$10 / MTok; Cache Hits & Refreshes=$0.50 / MTok; Output Tokens=$25 / MTok [source]
- Claude Opus 4.6: Base Input Tokens=$5 / MTok; 5m Cache Writes=$6.25 / MTok; 1h Cache Writes=$10 / MTok; Cache Hits & Refreshes=$0.50 / MTok; Output Tokens=$25 / MTok [source]
- Claude Opus 4.5: Base Input Tokens=$5 / MTok; 5m Cache Writes=$6.25 / MTok; 1h Cache Writes=$10 / MTok; Cache Hits & Refreshes=$0.50 / MTok; Output Tokens=$25 / MTok [source]
- Claude Opus 4.1 ([retired, except on Bedrock and Google Cloud](https://platform.claude.com/docs/en/about-claude/model-deprecations)): Base Input Tokens=$15 / MTok; 5m Cache Writes=$18.75 / MTok; 1h Cache Writes=$30 / MTok; Cache Hits & Refreshes=$1.50 / MTok; Output Tokens=$75 / MTok [source]
- Claude Opus 4 ([retired, except on Google Cloud](https://platform.claude.com/docs/en/about-claude/model-deprecations)): Base Input Tokens=$15 / MTok; 5m Cache Writes=$18.75 / MTok; 1h Cache Writes=$30 / MTok; Cache Hits & Refreshes=$1.50 / MTok; Output Tokens=$75 / MTok [source]
- Claude Sonnet 5: Base Input Tokens=$2 / MTok; 5m Cache Writes=$2.50 / MTok; 1h Cache Writes=$4 / MTok; Cache Hits & Refreshes=$0.20 / MTok; Output Tokens=$10 / MTok [source]
- Claude Sonnet 4.6: Base Input Tokens=$3 / MTok; 5m Cache Writes=$3.75 / MTok; 1h Cache Writes=$6 / MTok; Cache Hits & Refreshes=$0.30 / MTok; Output Tokens=$15 / MTok [source]
- Claude Sonnet 4.5: Base Input Tokens=$3 / MTok; 5m Cache Writes=$3.75 / MTok; 1h Cache Writes=$6 / MTok; Cache Hits & Refreshes=$0.30 / MTok; Output Tokens=$15 / MTok [source]
- Claude Sonnet 4 ([retired, except on Bedrock and Google Cloud](https://platform.claude.com/docs/en/about-claude/model-deprecations)): Base Input Tokens=$3 / MTok; 5m Cache Writes=$3.75 / MTok; 1h Cache Writes=$6 / MTok; Cache Hits & Refreshes=$0.30 / MTok; Output Tokens=$15 / MTok [source]
- Claude Haiku 4.5: Base Input Tokens=$1 / MTok; 5m Cache Writes=$1.25 / MTok; 1h Cache Writes=$2 / MTok; Cache Hits & Refreshes=$0.10 / MTok; Output Tokens=$5 / MTok [source]
- Claude Haiku 3.5 ([retired, except on Bedrock and Google Cloud](https://platform.claude.com/docs/en/about-claude/model-deprecations)): Base Input Tokens=$0.80 / MTok; 5m Cache Writes=$1 / MTok; 1h Cache Writes=$1.60 / MTok; Cache Hits & Refreshes=$0.08 / MTok; Output Tokens=$4 / MTok [source]
- TTL support — By default, automatic caching uses a 5-minute TTL. You can specify a 1-hour TTL at 2x the base input token price: [source]
- Pricing Changes for Cached Requests:: Cache write cost = Input token price + (Cache storage price × (5 minutes / 60 minutes)) — `Cache write cost = Input token price + (Cache storage price × (5 minutes / 60 minutes))` [source] — Gemini cache pricing formula as stated by OpenRouter
- Cache-Write Billing — For GPT-5.6 and later, OpenAI bills cache writes at 1.25x the model's uncached input rate — even with automatic caching, no opt-in required. Cache reads keep the discounted cache-read rate. [source] — OpenAI GPT-5.6 via OpenRouter
- Minimum token requirements — Each model has a minimum cacheable prompt length (see [Anthropic's cache limitations](https://platform.claude.com/docs/en/build-with-claude/prompt-caching#cache-limitations)): [source]
- [5m cache write](https://platform.claude.com/docs/en/build-with-claude/prompt-caching): $12.50 / MTok [source]
- [1h cache write](https://platform.claude.com/docs/en/build-with-claude/prompt-caching): $20 / MTok [source]
- [Cache read](https://platform.claude.com/docs/en/build-with-claude/prompt-caching): $1 / MTok [source]
- wrap: total_input_tokens = cache_read_input_tokens + cache_creation_input_tokens + input_tokens — `total_input_tokens = cache_read_input_tokens + cache_creation_input_tokens + input_tokens` [source]
- Batch requests are typically billed at 50% of the model's standard per-token pricing, though web-search calls bill at standard rates and prompt-caching rates vary. [source]
- Batch work that can wait — The [Batch API](https://platform.claude.com/docs/en/build-with-claude/batch-processing) takes 50% off every token of a request, including cached ones, in exchange for results arriving any time within 24 hours. [source]
- Prompt caching: Saving in these runs=Cost cut by a factor of 2.5 to 3.7 on agent loops; 83% on the triage run; Quality cost=None; Latency=Faster; Where=[Cache repeated context](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#cache-repeated-context) [source]
- 1-hour cache duration: Saving in these runs=Cheaper than the 5-minute default once about 1 turn in 20 follows a pause between 5 minutes and an hour and few gaps run over an hour; with no pauses the default cost 15% less on Claude Sonnet 5 and 11% less on Claude Opus 5; Quality cost=None; Latency=Stays warm after a pause; Where=[Pick the cache… [source]
Problems, failure modes and limitations
- Automatic prompt caching using the top-level `cache_control` field is not supported; explicit cache breakpoints must be used instead. [source] — scope: Claude on Amazon Bedrock (legacy) page only — on the Claude API the top-level cache_control IS supported (see Automatic caching)
- Prompt caching considerations — If you use [Prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching), changing the Skills list in your container breaks the cache. Skills render into the system prompt in a fixed order, so the same list produces the same cacheable prefix: [source]
- **Tool definitions**: Tools cache=✘; System cache=✘; Messages cache=✘; Impact=Modifying tool definitions (names, descriptions, parameters) invalidates the entire cache [source]
- **Tool choice**: Tools cache=✓; System cache=✓; Messages cache=✘; Impact=Changes to `tool_choice` parameter only affect message blocks [source]
- **Images**: Tools cache=✓; System cache=✓; Messages cache=✘; Impact=Adding/removing images anywhere in the prompt affects message blocks [source]
- Prompt caching is supported in the Message Batches API, but cache hits are provided on a best-effort basis due to concurrent and unordered processing. [source]
- [Web search](https://platform.claude.com/docs/en/agents-and-tools/tool-use/web-search-tool): Enabling or disabling invalidates the system and messages caches [source]
- [Web fetch](https://platform.claude.com/docs/en/agents-and-tools/tool-use/web-fetch-tool): Enabling or disabling invalidates the system and messages caches [source]
- Mid-conversation tool changes — The `tools` array sits even earlier in the hashed request prefix than the top-level `system` field, so editing it invalidates the [prompt cache](https://platform.claude.com/docs/en/build-with-claude/prompt-caching) for the entire conversation. [source]
- **Non-tool results passed to extended thinking requests**: Tools cache=✓; System cache=✓; Messages cache=Model-specific; Impact=On Opus 4.5+ and Sonnet 4.6+, thinking blocks are preserved by default, so the cache remains valid (✓). On earlier Opus/Sonnet models and all Haiku models, all previously-cached thinking blocks are stripped from context, and any messages that follow those thinking… [source]
- Setting `max_tokens` to 0 is not supported inside a batch because ephemeral cache entries written during processing likely expire before follow-up requests run. [source]
- **Web search toggle**: Tools cache=✓; System cache=✘; Messages cache=✘; Impact=Enabling/disabling web search modifies the system prompt [source]
- **Citations toggle**: Tools cache=✓; System cache=✘; Messages cache=✘; Impact=Enabling/disabling citations modifies the system prompt [source]
- **Speed setting**: Tools cache=✓; System cache=✘; Messages cache=✘; Impact=Switching between [`speed: "fast"` and standard speed](https://platform.claude.com/docs/en/build-with-claude/fast-mode) invalidates system and message caches [source]
- **Thinking parameters**: Tools cache=Model-specific; System cache=Model-specific; Messages cache=✘; Impact=The thinking configuration (mode, and `budget_tokens` in extended mode) is rendered into the prompt, so changing it always invalidates message blocks; tool and system caches are also invalidated on models that render the configuration ahead of them. See [Thinking and prompt… [source]
- **Effort setting**: Tools cache=Model-specific; System cache=Model-specific; Messages cache=✘; Impact=Changing the [`output_config.effort`](https://platform.claude.com/docs/en/build-with-claude/effort) value always invalidates message blocks, with the same model-specific effect on tool and system caches as thinking parameters. Setting effort explicitly to the model's default is equivalent to… [source]
- Limitations — A `max_tokens: 0` request is rejected with an `invalid_request_error` if any of the following are set, since each implies output that a zero-token budget cannot produce: [source]
- What invalidates the cache — Modifications to cached content can invalidate some or all of the cache. [source]
- Cache hits drop after changing thinking settings — `cache_read_input_tokens` falls to zero on requests that previously hit the cache. [source]
- Modifying tool definitions: Entire cache (tools, system, messages) [source]
- Toggling web search or citations: System and messages caches [source]
- Changing `tool_choice`: Messages cache [source]
- Changing `disable_parallel_tool_use`: Messages cache [source]
- Toggling images present/absent: Messages cache [source]
- Changing thinking parameters: Messages cache always; tool and system caches too on models that render the thinking configuration ahead of them ([details](https://platform.claude.com/docs/en/build-with-claude/thinking#thinking-and-prompt-caching)) [source]
- Interaction with Prompt Caching — Auto Exacto reorders providers on every tool-calling request, which can conflict with the sticky routing used by [prompt caching](/docs/guides/best-practices/prompt-caching). [source]
- [Computer use](https://platform.claude.com/docs/en/agents-and-tools/tool-use/computer-use-tool): Screenshot presence affects messages cache; `cache_control` goes on the toolset entry (see [cache\_control on tool definitions](https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-use-with-prompt-caching#cache-control-on-tool-definitions)) [source]
- [Browser use](https://platform.claude.com/docs/en/agents-and-tools/tool-use/browser-use-tool): Screenshot presence affects messages cache; `cache_control` goes on the toolset entry (see [cache\_control on tool definitions](https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-use-with-prompt-caching#cache-control-on-tool-definitions)) [source]
- Why am I seeing 'TypeError: Cannot read properties of undefined (reading 'messages')'?: client.messages.create(/* ... */); — `client.messages.create(/* ... */);` [source]
- Cache limitations — On the Claude API, [Claude Platform on AWS](https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws), [Google Cloud](https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai), and [Microsoft… [source]
- Changing `output_config.effort`: Same as thinking parameters; setting the model's default explicitly is equivalent to omitting it [source]
Comparisons and alternatives
- What stays the same — Automatic caching uses the same underlying caching infrastructure. Pricing, minimum token thresholds, context ordering requirements, and the 20-block lookback window all apply the same as with explicit breakpoints. [source]
- Explicit Prompt Caching — Automatic prompt caching still works with no code changes. Explicit caching adds direct control over cache boundaries instead of relying on OpenAI's automatic breakpoint placement. [source]
- Prompt caching: What it reduces=Token cost of repeated tool definitions; When it fits=Stable toolsets across many requests; Learn more=[Tool use with prompt caching](https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-use-with-prompt-caching) [source]
Changes and history
- Replacing the max\_tokens=1 workaround — Before `max_tokens: 0` was available, some applications used `max_tokens: 1` warm-up calls to achieve the same effect. [source]
- GPT-5.6 Migration Guide — Adopt reasoning.mode, reasoning.context, and explicit prompt caching for the GPT-5.6 model family [source]
- Python: client.beta.prompt_caching.messages.create(**params) — `client.beta.prompt_caching.messages.create(**params)` [source] — deprecated beta namespace, quoted from a TypeError FAQ; current SDKs use client.messages.create with cache_control on blocks
- TypeScript: client.beta.promptCaching.messages.create(/* ... */); — `client.beta.promptCaching.messages.create(/* ... */);` [source] — deprecated beta namespace, quoted from a TypeError FAQ; current SDKs use client.messages.create with cache_control on blocks
Facts and statements
- [Code execution](https://platform.claude.com/docs/en/agents-and-tools/tool-use/code-execution-tool): Container state is independent of prompt cache [source]
- Grouping requests across modalities — Beyond sticky routing, OpenRouter uses `session_id` to group your requests in the [Sessions view on the Logs page](https://openrouter.ai/logs?tab=sessions). One `session_id` links requests across conversation turns, retries, and different modalities. [source]
- [Fallback credit](https://platform.claude.com/docs/en/build-with-claude/fallback-credit): Description=Avoid paying the prompt-cache cost twice when you retry a refused request on another model. The refusal carries a credit token, and echoing it on the retry bills the retry as though the conversation had been on the new model all along. Message Batches results do not include fallback credit… [source] — fallback credit: retrying a refused request without paying the cache twice
- Supported Models and Limitations: — Only certain Gemini models support caching. Please consult Google's [Gemini API Pricing Documentation](https://ai.google.dev/gemini-api/docs/pricing) for the most current details. [source]
- [Text editor](https://platform.claude.com/docs/en/agents-and-tools/tool-use/text-editor-tool): Standard client tool, no special caching interaction [source]
- [Bash](https://platform.claude.com/docs/en/agents-and-tools/tool-use/bash-tool): Standard client tool, no special caching interaction [source]
- [Memory](https://platform.claude.com/docs/en/agents-and-tools/tool-use/memory-tool): Standard client tool, no special caching interaction [source]
- Feature compatibility — Citations work in conjunction with other API features including [prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching), [token counting](https://platform.claude.com/docs/en/build-with-claude/token-counting), and [batch… [source]
- Supported models — Prompt caching (both automatic and explicit) is supported on all [active Claude models](https://platform.claude.com/docs/en/models/overview). [source]
- Data retention — Prompt caching (both automatic and explicit) is ZDR eligible. Anthropic does not store the raw text of your prompts or Claude's responses. [source]
- Mechanics — Three mechanics follow from Claude managing its own thinking: turn validation, prompt caching, and how you bound cost. [source]
- Prompt caching — For how `defer_loading` preserves prompt caching, see [Tool use with prompt caching](https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-use-with-prompt-caching). [source]
- Prompt caching — For caching tool definitions across turns, see [Tool use with prompt caching](https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-use-with-prompt-caching). [source]
- Billing — You are not billed for a request that is refused before any output is generated. When you retry on another model, [fallback credit](https://platform.claude.com/docs/en/build-with-claude/fallback-credit) refunds the prompt-cache cost of switching, so you avoid paying that cost twice. [source] — fallback credit: retrying a refused request without paying the cache twice
- Mid-Conversation Tool Changes (Beta) — Opus 5 supports [changing the available tool set between turns](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages#mid-conversation-tool-changes) without invalidating the prompt cache, using system-role content blocks: [source]
- The top-level `system` field never changes in this example to ensure the cached prefix stays intact. [source]
- Data retention — Cache diagnostics is ZDR eligible (qualified). Anthropic does not store the raw text of your prompts or Claude's outputs for this feature. [source]
- Must match exactly: `system`, `messages`, `tools`, `tool_choice`, `thinking`, and `cache_control`, plus `output_config`, `mcp_servers`, `context_management`, and `container` when you use them [source] — fallback credit: retrying a refused request without paying the cache twice
- Checking that the credit applied — The refund is visible in the retry's `usage`. Compared with what the same request would report without the token, `cache_creation_input_tokens` is lower, and `cache_read_input_tokens` is higher by the same amount. [source] — fallback credit: retrying a refused request without paying the cache twice
- `usage.prompt_tokens_details`: Always empty [source] — scope: the OpenAI SDK compatibility layer of the Claude API — the field is always empty there; OpenRouter does populate prompt_tokens_details
Related concepts
- cache_control — is a part of Prompt caching; Anthropic/OpenRouter per-block cache marker
- cache breakpoint — is a part of Prompt caching
- cache hit — is a measure of Prompt caching
- cache pricing — is a measure of Prompt caching; pricing multipliers and table columns; MTok alone is any price table
- prefix hierarchy — is a part of Prompt caching; tools → system → messages; a change at one level invalidates it and everything after
- cache write — is a measure of Prompt caching
- Message Batches — is a related of Prompt caching; features whose interaction with the prompt cache the docs describe; never qualifies a unit alone
- cache diagnostics — is a part of Prompt caching; Anthropic beta that reports why a cache missed
- automatic caching — is a hyponym of Prompt caching; provider places the cache point (Anthropic auto mode, OpenAI, Gemini implicit)
- cache TTL — is a measure of Prompt caching
- cache invalidation — is a problem of Prompt caching; what changes cause the prefix to stop matching
- cache prefix — is a part of Prompt caching; the cache key is the exact request prefix: tools → system → messages
- cache_read_input_tokens — is a measure of Prompt caching; usage/pricing fields reported per response
- cacheable — is a variant of Prompt caching; the docs' adjective for content the prompt cache can hold
- cache read — is a measure of Prompt caching
- explicit caching — is a hyponym of Prompt caching; caller places cache_control breakpoints
- cache pre-warming — is a dependent of Prompt caching; loading the cache before real traffic
- sticky routing — is a dependent of Prompt caching; OpenRouter: routing follow-up requests to the provider holding the cache
- ephemeral cache_control type — is a instance of Prompt caching; the only cache_control type; bare 'ephemeral' excluded (containers)
- minimum cacheable prompt length — is a measure of Prompt caching; per-model token floors below which nothing is cached
Children
- No children recorded.