Prompt caching reuses the computation for the part of a request that is identical to the start of an earlier request (the prefix), so that part of the input is processed more cheaply and more quickly. Both the OpenAI API and the Claude API offer it, but how you turn it on, how long the cache lives, what a cache write costs, and how you check whether it is working differ a lot between the two. If you design for both with the same assumptions, you can end up with caching that works on one side while, on the other, you pay the write price on every request and never get a hit.

This article compares OpenAI and Anthropic prompt caching based on the original text of OpenAI's "Prompt caching," "Prompt cache diagnostics," and "Pricing" docs and Anthropic's "Prompt caching," "Cache diagnostics," "Pricing," and "Rate limits" docs. It covers how the two specs differ, how to get cache hits, and how to confirm you are getting them. Every number was checked against the original pages on October 3, 2026. For the bigger picture on cutting API costs (choosing models, batching, controlling output, and so on), see our separate article "How to cut your AI usage costs."

The short answer: 4 ways the two caches differ

Sources: OpenAI "Prompt caching," Anthropic "Prompt caching" (checked October 3, 2026)

How to turn it on

OpenAI: on by default

Claude caches only when the request includes cache_control.

TTL

30 min vs. 5 min / 1 hour

OpenAI (GPT-5.6 and later): at least 30 minutes after last use. Claude: 5 minutes by default, with a 1-hour option.

Pricing

Cache writes cost extra

Both charge 1.25x the input price for writes (2x for Claude's 1-hour cache). Reads are usually 0.1x, and cheaper still on some models.

How to check

usage and diagnostics

input_tokens means different things on each side. Both offer a diagnostics feature that compares a request with an earlier one to explain a miss.

1. What prompt caching is: reusing an identical prefix

Every time a language model reads its input, it computes intermediate values for each token (the KV, or key and value, tensors). Prompt caching stores those values from the start of the prompt up to a certain point and skips the computation when the next request begins with exactly the same tokens. OpenAI's guide explains that what is stored is the KV values, not the tokens themselves.

The key point is that only the part that matches from the very start can be reused. In the two requests below, only the instructions and documents can be reused.

Request 1: [Instructions 5,000 tokens][Documents 20,000 tokens][Question A]
Request 2: [Instructions 5,000 tokens][Documents 20,000 tokens][Question B]
           └────────────── identical up to here ──────────────┘└ differs  ┘
→ The 25,000 tokens of instructions + documents can be reused

Request 3: [Today's date][Instructions 5,000 tokens][Documents 20,000 tokens][Question C]
           └─ differs ──┘
→ Even though the rest matches, not a single token can be reused

Both vendors' docs state that caching does not change the content of the output. It does not store and replay a previous answer; it only skips the work of reading the input. That is why it still helps in chats where every question is different, or in agents that read different documents each time, as long as there is a shared prefix (instructions, tool definitions, conversation history).

2. OpenAI vs. Anthropic prompt caching, side by side

On September 22, 2026, OpenAI announced caching improvements for GPT-6, and the mechanism changed for GPT-5.6 and later models (30-minute retention, explicit breakpoints, paid cache writes, and more). The OpenAI column in the table below describes GPT-5.6 and later. The differences for GPT-5.5 and earlier follow the table.

ItemOpenAI (GPT-5.6 and later)Claude (Anthropic)
How to turn it onOn by default for supported models; prompt_cache_options.mode chooses implicit or explicit-onlyOnly with cache_control (one at the top level for automatic caching, or on individual blocks as explicit breakpoints)
Number of breakpointsUp to 4 cache writes per requestUp to 4
TTL (retention)At least 30 minutes after the last write or reuse (ttl accepts only "30m")5 minutes by default, 1 hour with "ttl": "1h"; both refresh each time the cache is used
Cache write price1.25x input1.25x input for 5 minutes, 2x for 1 hour
Cache read price0.1x input (0.05x for GPT-6.1 Sol)0.1x input (0.05x for Opus 5.5, 0.025x for Fable 5.1 and Mythos 5.1)
Minimum length1,024 tokens of visible input512 to 4,096 tokens depending on the model (table in section 3)
Sharing scopePer organization (not shared across processing regions)Per workspace on the Claude API (per organization on Bedrock and Google Cloud)
Rate limitsTokens read from the cache still count toward TPMOn most models, tokens read from the cache do not count toward the input limit (ITPM)
Prewarmingprompt_cache_options.prewarm: trueSend with max_tokens: 0
Usage fieldscached_tokens, cache_write_tokenscache_read_input_tokens, cache_creation_input_tokens
Miss diagnosticscomparison_response_id → prompt_cache_diagnostics (Responses API)diagnostics.previous_message_id → diagnostics (Claude API only)

Sources: OpenAI "Prompt caching," "Prompt cache diagnostics"; Anthropic "Prompt caching," "Cache diagnostics," "Rate limits" (checked October 3, 2026)

GPT-5.5 and earlier models only have implicit caching, with breakpoints placed automatically at fixed intervals, and no extra charge for cache writes. Retention is set with prompt_cache_retention: according to the guide, in_memory lasts "about 5 to 10 minutes of inactivity, up to 1 hour," and 24h lasts "usually about 30 minutes, up to 24 hours." When you move to GPT-5.6 or later, replace this setting with prompt_cache_options.ttl.

Per-token prices for each model are covered in our pricing comparison, "Claude vs. ChatGPT pricing." For the GPT-6 models (Astra, Sol, Luna), see "our GPT-6 Sol and Luna article," and for every vendor's current models, see "AI model knowledge cutoff dates."

3. When the cache hits: prefix, minimum length, TTL, and scope

The shared prefix and the order of the request

On both platforms, the cache hits only when the prefix up to the breakpoint matches exactly. Claude reads the request from the start in the order tools → system → messages, so changing a single tool definition invalidates the cache for the system prompt and the conversation history that follow it. OpenAI likewise explains that tool definitions, the output format (text.format), reasoning effort (reasoning.effort), and similar settings are part of the prefix.

In practice, the layout is the same for both: put what does not change (tool definitions, instructions, documents) first, and what changes on every request (dates, per-user data, the question) last. Add to a conversation by appending, without rewriting the history.

Placing breakpoints: automatic or manual

OpenAI's implicit mode (GPT-5.6 and later) places a breakpoint at the end of the most recent eligible message (a user message, the last of a run of tool results, and so on). In explicit-only mode, only the positions where you add prompt_cache_breakpoint become breakpoints, and if you add none, the cache is not used and you pay no write price either.

Claude's automatic caching, enabled with a single top-level "cache_control": {"type": "ephemeral"}, places a breakpoint on the last cacheable block and moves it forward as the conversation grows. Adding cache_control to individual blocks lets you choose the breakpoints yourself.

// Claude: breakpoint at the end of the unchanging system prompt (explicit breakpoint)
{
  "model": "claude-sonnet-5-5",
  "max_tokens": 1024,
  "system": [
    {
      "type": "text",
      "text": "Long instructions and documents...",
      "cache_control": { "type": "ephemeral" }
    }
  ],
  "messages": [{ "role": "user", "content": "Today's question..." }]
}

// OpenAI (Responses API): breakpoint after the unchanging instructions, explicit-only mode
{
  "model": "gpt-6.1-sol",
  "prompt_cache_options": { "mode": "explicit" },
  "input": [
    {
      "role": "developer",
      "content": [{
        "type": "input_text",
        "text": "Long instructions and documents...",
        "prompt_cache_breakpoint": { "mode": "explicit" }
      }]
    },
    { "role": "user", "content": "Today's question..." }
  ]
}

Minimum length: short prefixes are not cached

For GPT-5.6 and later, OpenAI's minimum is 1,024 tokens of visible input (hidden instructions that OpenAI adds behind the scenes do not count). On Claude, it depends on the model.

Minimum lengthClaude models
512 tokensFable 5.1, Mythos 5.1, Opus 5.5, Opus 5, Sonnet 5.5, Fable 5, Mythos 5
1,024 tokensOpus 4.8, Sonnet 5, Sonnet 4.6, Sonnet 4.5, and others
2,048 tokensOpus 4.7, Mythos Preview
4,096 tokensOpus 4.6, Opus 4.5, Haiku 4.5

Source: Anthropic "Prompt caching," Cache limitations (checked October 3, 2026). On Bedrock, the values in AWS's documentation apply.

If a Claude prompt is below the minimum, adding cache_control does not raise an error; the prompt is silently not cached. In that case, cache_creation_input_tokens and cache_read_input_tokens in usage are both 0. The minimum changes when you switch models, so a prefix that was cached on the previous model may stop being cached (OpenAI's guide gives the same warning).

TTL: watch where the clock starts

For OpenAI (GPT-5.6 and later), the cache lasts "at least 30 minutes after the last write or reuse, and may persist longer." Claude defaults to 5 minutes, and choosing 1 hour makes writes cost 2x the input price. On both platforms, the TTL is refreshed at no extra cost each time the cache is used.

Claude has one trap: the TTL is counted from the start of the request, not from the end of the response. In the docs' example, if a response takes 4 minutes to generate, the next request must start within about 1 minute of that response finishing to hit the 5-minute cache. For agents that generate long outputs, 5 minutes is shorter than it looks.

Cache scope and where the cache lives

OpenAI's cache is per organization and is not shared across processing regions (data residency settings). The guide also says the cache lives on individual machines, and above roughly 15 requests per minute, requests may be routed to a different machine. From GPT-5.6 onward, OpenAI handles routing automatically, and prompt_cache_key has become an optional setting for splitting cache reporting by customer rather than a way to raise the hit rate (on GPT-5.5 and earlier, using the same key to steer requests to the same machine mattered).

The Claude API scopes the cache per workspace. Within the same organization, different workspaces do not share a cache even for identical prompts. On Bedrock and Google Cloud, it is per organization. Also, the cache becomes available only once the first response begins, so if you send many requests with the same prefix in parallel all at once, the others become writes too before the first one has written the cache.

4. Pricing and break-even: how many reads before caching pays off

Because a cache write costs more than regular input, a cache entry that is written and never read costs more than not caching at all. Let w be the write multiplier, r the read multiplier, and n the number of reads after the write. The number of reads needed to break even follows from these formulas.

With caching    = w + n × r
Without caching = 1 + n          (sending the same prefix n + 1 times as is)
Pays off when   : n > (w − 1) ÷ (1 − r)
SettingWrite wRead rAfter 1 read (no cache = 2)Reads to break even
OpenAI, most GPT-5.6+ models1.250.11.351
OpenAI GPT-6.1 Sol1.250.051.301
OpenAI GPT-5.5 and earlierNo write chargeVaries by model—Never a loss
Claude 5-minute (most models)1.250.11.351
Claude 1-hour (most models)20.12.10 (loss)2 (2.20 vs. 3)
Claude 1-hour (Opus 5.5)20.052.05 (loss)2 (2.10 vs. 3)
Claude 1-hour (Fable 5.1)20.0252.025 (loss)2 (2.05 vs. 3)

Multipliers relative to an input price of 1. Calculated from OpenAI "Prompt caching" and "Pricing" and Anthropic "Pricing" (checked October 3, 2026). Compares the prefix only; output and the per-request question are excluded.

OpenAI's guide includes the same calculation for a 0.1x model: "writing once and reading once costs 1.35x, versus 2x for processing it twice without caching." Anthropic's pricing page likewise says the 5-minute cache pays off after one read and the 1-hour cache after two. No matter how cheap reads are, a 1-hour cache cannot pay off with a single read, because the 2x write is too heavy.

Comparing models with the same input price: request spacing flips the result

At official prices as of October 3, 2026, GPT-6.1 Sol and Claude Sonnet 5.5 both charge $2 per million input tokens, and their 5-minute write price is the same $2.50. What differs is the read price (Sol $0.10, Sonnet 5.5 $0.20) and the TTL. We calculated the cost of the prefix for sending a 100,000-token prefix 10 times at different intervals.

Request intervalNo cacheGPT-6.1 SolSonnet 5.5 (5-minute)Sonnet 5.5 (1-hour)
Every 3 minutes$2.00$0.34$0.43$0.58
Every 20 minutes$2.00$0.34$2.50 (write every time)$0.58
Every 45 minutes$2.00Up to $2.50 (no guarantee past 30 minutes)$2.50$0.58
Every 2 hours$2.00Up to $2.50$2.50$4.00

Prices from OpenAI "Pricing" (Standard, up to 272K) and Anthropic "Pricing" (both checked October 3, 2026). 100,000 tokens = 0.1 (in units of 1 million tokens). On hits, the first request is a write and the other 9 are reads; on misses, all 10 are writes. OpenAI can also miss because of machine routing and similar factors, so the hit rows are values under favorable conditions.

For example, GPT-6.1 Sol every 3 minutes works out to "0.1 × $2.50 (one write) + 0.1 × $0.10 × 9 reads = $0.25 + $0.09 = $0.34." The table shows three things.

  • At intervals of 5 minutes or less, the gap between the two is small ($0.34 vs. $0.43). It comes down to the read price alone.
  • At intervals of 5 to 30 minutes, OpenAI's 30-minute TTL wins. With Claude's 5-minute TTL, every request becomes a write, costing $2.50, which is more than no cache at all. Switching to 1 hour brings it down to $0.58.
  • When the interval exceeds the TTL, caching loses money. At every 2 hours, no cache at $2.00 is the cheapest. On Claude, leave off cache_control; on OpenAI GPT-5.6 and later, use explicit-only mode and place no breakpoints, and you avoid the write charge.

In particular, because caching is on by default for OpenAI GPT-5.6 and later, staying in implicit mode can mean paying the write price even for input you will never send again. If your workload sends many long, one-off inputs, check cache_write_tokens in usage and decide whether to switch to explicit-only mode.

How caching combines with other pricing

  • Batch: OpenAI's price list also has cached input and cache write prices for Batch and Flex (GPT-6.1 Sol on Batch: $1 input, $0.05 cache reads, $1.25 cache writes). Anthropic says the cache multipliers stack with the 50% batch discount, but because batch requests are processed in parallel in no particular order, it describes cache hits as "best effort."
  • Long inputs: On OpenAI, when input exceeds 272K tokens, the input, cache read, and cache write prices each double (the multipliers stay the same). On Anthropic, Claude 4.6 and later models charge the same price up to 1 million tokens.
  • Prewarming: On both, a prewarm write is billed at the normal write price. Claude's max_tokens: 0 incurs no output charge.

5. Why prompt caching is not working: common causes

Combining both vendors' notes on common pitfalls with the lists of reasons their diagnostics return, the causes of cache misses fall into roughly three groups.

The prefix changed

The instructions include a date or request ID. Tools are in a different order each time. The history was summarized, trimmed, or reordered. The JSON is built in a language whose key order changes between runs.

A setting changed

The model was switched (fallback, A/B test), or the reasoning effort, output format, Claude's thinking settings or the presence of images, or OpenAI's service tier differs from the previous request.

Conditions not met

The prefix is below the minimum length. The TTL expired. The requests were sent in parallel all at once. On Claude, the request came from a different workspace.

Common issues with OpenAI

  • No breakpoint right after the shared prefix: Implicit mode places the breakpoint at the end of the latest message, so with "fixed instructions + a different user message each time," the varying part gets written too and the next request misses. Put an explicit breakpoint right after the fixed part.
  • Switching to explicit-only mode midway: Explicit-only mode looks only for breakpoints you placed, so it does not hit cache entries written in implicit mode.
  • Appending to the same message: If a message that ended with "content A" becomes "content A + content B," the earlier breakpoint now falls in the middle of a message and misses. Add the new content as a new message instead.
  • Changing reasoning effort midway: On GPT-6 models, you can change the effort without breaking the cache by leaving the request's reasoning.effort as is and appending a configuration_update after the input.
  • Running compaction (context compression): The prefix changes, so the hit rate drops. The guide notes, however, that total cost can still go down because the input shrinks, and advises comparing total cost.

Common issues with Claude

  • A breakpoint on a block that changes every time: Writes happen only at breakpoint positions, and reads only look backward for earlier write positions. If a breakpoint sits on a block that changes every time, you pay the write price every time and never get a hit. Automatic caching also puts its breakpoint on the last block, so it falls into the same trap. Put an explicit breakpoint on the last block that does not change.
  • 20 or more blocks added in one turn: The lookup for earlier writes covers up to 20 positions back from a breakpoint. If the conversation grows by a lot at once, the earlier write falls outside that window. Keep an extra breakpoint earlier in the prompt.
  • Rewriting the system prompt midway: On supported models, you can add instructions without breaking the cache by leaving the top-level system unchanged and adding a message with "role": "system" inside messages.
  • Switching between fast mode (speed: "fast") and standard: This invalidates the system and conversation caches.

6. How to check cache hits: usage and diagnostics

Start with usage: input_tokens means different things

On both platforms, the response's usage reports how much was read from and written to the cache. The thing to watch is that input_tokens means opposite things on the two platforms.

What you want to knowOpenAI (Responses API)Claude
Tokens read from cacheusage.input_tokens_details.cached_tokensusage.cache_read_input_tokens
Tokens written to cacheusage.input_tokens_details.cache_write_tokensusage.cache_creation_input_tokens (5-minute vs. 1-hour breakdown in cache_creation)
What input_tokens containsTotal input (including reads and writes)Only the tokens after the last breakpoint, unrelated to the cache
Total inputinput_tokenscache_read_input_tokens + cache_creation_input_tokens + input_tokens
Hit ratecached_tokens ÷ input_tokenscache_read_input_tokens ÷ the total above

Sources: OpenAI "Prompt caching," Monitor cache performance; Anthropic "Prompt caching," Tracking cache performance (checked October 3, 2026)

If you divide by Claude's input_tokens as if it were total input, both your hit rate and your cost will be badly off. When you put both vendors' numbers on the same dashboard, normalize the totals with the formulas above before comparing. OpenAI's guide recommends calculating "token hit rate" as total tokens read from the cache divided by total input tokens, aggregated by user, by day, and so on.

For cost, OpenAI is "(input − reads − writes) × price + reads × price × 0.1 + writes × price × 1.25," and Claude is "input_tokens × price + reads × price × 0.1 + writes × price × 1.25 (2 for the 1-hour portion)" (replace the read multiplier with 0.05 or another value depending on the model). OpenAI has a "Prompt Caching Dashboard" on its usage page, and Anthropic's Rate limits documentation points you to the Usage page to see your cache hit rate.

What to check, in order, when reads are 0

  1. Are writes also 0? On Claude, if both are 0, the prompt is below the minimum length or cache_control is missing. On OpenAI in explicit-only mode, no write happens if you placed no breakpoint either.
  2. Do writes appear on every request? The breakpoint sits on a position that changes each time, or something in the prefix changes each time. Use the diagnostics described next to find what changed.
  3. How long since the previous request? Check whether it exceeded Claude's 5 minutes (including response generation time) or OpenAI's 30 minutes.

OpenAI diagnostics: prompt_cache_diagnostics

In OpenAI's Responses API, on supported GPT-5.6 and later models, passing a previous response's ID in prompt_cache_options.comparison_response_id puts the result of comparing that request with the current one into prompt_cache_diagnostics. The docs say this costs nothing extra and does not count separately against rate limits.

// Add a comparison target to the second request
{
  "model": "gpt-6.1-sol",
  "input": [ ...same prefix as the first request..., { "role": "user", "content": "Next question" } ],
  "prompt_cache_options": { "comparison_response_id": "resp_(ID of the first response)" }
}

// Example returned on a miss (the docs' example: a tool was renamed)
{
  "prompt_cache_diagnostics": {
    "type": "cache_miss",
    "reason": "tools_changed",
    "comparison_reusable_tokens": 5629,
    "cache_missed_tokens": 5629
  }
}

There are four type values: cache_hit, cache_miss, comparison_response_not_found (no comparison record, or it expired), and unavailable (no conclusion). On a miss, reason is one of these nine.

reasonWhat changed
model_changedA different model handled the request (routing, A/B test, fallback)
prompt_cache_key_changedprompt_cache_key changed (may be counted as a miss even when the cache actually still exists)
service_tier_changedThe service tier changed (the request may also be processed on a tier other than the one specified)
tools_changedTools were added, removed, or reordered, or their descriptions or schemas changed
text_format_changedThe output format or its schema changed
reasoning_effort_changedThe reasoning effort changed
verbosity_changedThe response verbosity changed
context_compactedCompaction replaced the earlier conversation
input_changedEarlier input changed (a timestamp or ID in the instructions, or history edited, reordered, or deleted)

Source: OpenAI "Prompt cache diagnostics," Fix a cache miss (checked October 3, 2026)

Claude diagnostics: diagnostics

On the Claude API, you need to include a diagnostics field in every request, because the API stores a comparison fingerprint (hashes and estimated token counts) only for requests that include it. On the first turn, pass "previous_message_id": null; from then on, pass the previous response's id.

// From the second turn onward
{
  "model": "claude-sonnet-5-5",
  "max_tokens": 1024,
  "cache_control": { "type": "ephemeral" },
  "diagnostics": { "previous_message_id": "msg_(id of the previous response)" },
  "system": "...",
  "messages": [ ... ]
}

If the response's diagnostics is null, no difference was found (or no comparison was made); {"cache_miss_reason": null} means the comparison has not finished yet; and if a reason is present, it marks the first place where the requests diverged. There are six reasons: model_changed, system_changed, tools_changed, messages_changed, previous_message_not_found, and unavailable. The *_changed reasons come with cache_missed_input_tokens, an estimate of how much was lost.

Claude's documentation says to read the diagnostics together with cache_read_input_tokens.

Diagnostics resultCache readsWhat it means
nullHighHitting as expected
nullLow or 0The request is the same, but the cache had expired (shorten the interval or use 1 hour)
*_changedLow or 0The request changed (fix the location the reason points to)
*_changedHighA rare case: something changed further back, but an earlier breakpoint still hit

Source: Anthropic "Cache diagnostics," Reading diagnostics alongside usage (checked October 3, 2026)

The two diagnostics features have a lot in common: both return only the first difference found, so after fixing it, compare again. The differences are that Claude's diagnostics work only on the Claude API (not on Amazon Bedrock, Google Cloud, Claude Platform on AWS, or Microsoft Foundry), and only compare against requests from the same workspace. OpenAI's documentation does not describe any step for marking the earlier request that will be compared against. On Claude, if the earlier request did not include diagnostics as well, you get previous_message_not_found.

7. Cache hit rate figures: vendor examples and third-party data

The cache hit rate figures in circulation mean very different things depending on who published them and under what conditions. Here they are, separated.

Figures from the vendors

  • Examples in OpenAI's guide: A one-off grading workload (using an LLM as a grader) with an explicit breakpoint after a fixed rubric reached a "token hit rate of about 70%," and an agent that calls tools repeatedly reached "over 90%." The guide cautions that these are examples of possible results and that the ceiling depends on the workload.
  • Customer comments in OpenAI's announcement (September 22, 2026): The Manus team says that after rethinking breakpoint placement, its hit rate on OpenAI models went from "about 85% to consistently over 90%" in less than a week. A comment about GitHub Copilot says that over the past few months it has cut the share of input that must be reprocessed by more than 50% compared with its earlier baseline. Both are customer quotes that OpenAI published on its own page, not independent measurements.
  • Anthropic: The docs give no real-world hit rate figures. The Rate limits page has a worked example, "at an 80% hit rate, an input limit of 2 million tokens per minute effectively processes 10 million tokens per minute," but that is arithmetic on an assumption, not a measurement.

Third-party data: not a direct verdict on which is better

Requesty, a service that routes requests to multiple AI APIs, aggregated the requests that passed through its own gateway and published hit rates for April 2026 of 77% for Anthropic direct (77.50% in the table) and 36% for OpenAI (36.40%) ("Prompt-cache hit rate per provider, April 2026," updated May 9). At face value, Claude appears to hit more than twice as often, but there are four reasons these numbers cannot be used for the comparison in this article.

  1. The data predates OpenAI's change: OpenAI introduced 30-minute retention and explicit breakpoints on September 22, and the April data comes from before that.
  2. The workloads differ: These are the results of different users' apps sending different prompts at different intervals through the gateway, not the same prompt sent to both vendors. On Claude, only requests where users added cache_control themselves are counted.
  3. The denominator: The page says "cached_tokens ÷ input_tokens," but as explained in section 6, Claude's input_tokens excludes cached tokens. The page does not say how it converted them.
  4. The page contradicts itself: In one place it puts Claude via Google Cloud (Vertex) at 24%, and in another at 14%.

As of our search on October 3, 2026, we found no third-party measurement that compared the two caches under the same conditions. In the end, the only way to know whether caching works in your app is to measure it in your own usage data. Calculate the hit rate and cost with the formulas in section 6, and if you are missing, use the diagnostics to find out why.

Summary

OpenAI and Claude prompt caching share the same basic idea: reuse the identical prefix and bill reads at around 0.1x the input price. The difference is that OpenAI (GPT-5.6 and later) is on by default, keeps the cache for at least 30 minutes, and charges 1.25x for writes, while Claude caches only when you add cache_control, keeps the cache for 5 minutes (with a 1-hour option), and charges 1.25x for writes (2x for 1 hour).

On pricing, because writes cost extra, a cache that is never read is a loss. The 5-minute and 30-minute caches pay off after one read, and Claude's 1-hour cache after two. With request intervals of 5 to 30 minutes, OpenAI's 30-minute TTL has the edge, and on Claude, choosing 1 hour avoids the reversal. If your requests are spaced further apart than the TTL, not caching is cheaper.

Check whether caching works by looking at the read and write amounts in usage. Claude's input_tokens covers only what comes after the breakpoint, so compute the total first, then the hit rate. When you miss, OpenAI's prompt_cache_diagnostics and Claude's diagnostics tell you where the request differs from the earlier one. For other ways to cut costs, see "How to cut your AI usage costs."

FAQ

Q. What is the TTL for OpenAI prompt caching?

A. For GPT-5.6 and later, it is at least 30 minutes after the last write or reuse. The prompt_cache_options.ttl setting accepts only "30m", and the guide says the cache may persist longer. For GPT-5.5 and earlier, you choose with prompt_cache_retention: in_memory lasts about 5 to 10 minutes of inactivity (up to 1 hour), and 24h lasts up to 24 hours.

Q. Can I extend the TTL for Claude prompt caching?

A. Yes, "cache_control": {"type": "ephemeral", "ttl": "1h"} gives you 1 hour. Writes cost 2x the input price, so it does not pay off unless the cache is read at least twice. If you keep using it at intervals under 5 minutes, the 5-minute cache is refreshed for free on every read. Both are counted from the start of the request, so the time spent generating a long response counts toward the TTL.

Q. Does caching change the answers?

A. No. Both vendors' docs say caching does not affect output generation. What is stored is the intermediate computation from reading the input, not the answer itself. As without caching, the same input does not always produce the same answer.

Q. Can I clear the cache manually?

A. Neither platform lets you. OpenAI says entries expire according to the TTL and settings, and Anthropic says they are removed automatically after at least 5 minutes without use (1 hour if you chose 1 hour). If you want to replace the prompt's content, change the prefix and the next request will write a new entry.

Sources

All sources were checked in the original on October 3, 2026. Prices, TTLs, and minimum lengths can change as models are added, so confirm them on each vendor's pricing page before relying on them. The cost calculations in this article are arithmetic based on official prices, not measurements from calling the APIs ourselves.