Prompt Caching Is Not 10% Everywhere: What a Cache Hit Costs Across Eight Providers
A cache hit is the cheapest token you can buy: the part of a prompt the provider has already seen and does not re-process. What it costs is not the same everywhere, and neither is what it costs to put something into the cache in the first place. This post is a reading of the site’s pricing table, column “Cached”, as it stands on September 17, 2026: 49 of the 64 current models with a verified price publish a cache-hit price, from eight providers. Every figure below was checked by hand against the page each row cites, on September 15, or on September 16 and 17 for Meta and for the rows added to the dataset on those days; more on why “by hand” matters at the end.
Who publishes nothing: Cohere, for all four of its current priced models; Mistral for seven of its nine (the exceptions are the two GLM rows); and Google for four, Gemini Omni Flash, Gemini Omni Flash Preview, Gemini 2.5 Computer Use Preview and Gemini Robotics ER 2 Streaming Preview. Where a provider prints no cache price, the table shows none, and this post does not estimate one.
The norm: 10% of the input price
Thirty-one of the 49 rows sit at exactly 0.1×, and they come from four providers. OpenAI, all eight current models: GPT-6 Astra $1.00 on $10.00, GPT-5.6 Sol $0.40 on $4.00, Terra $0.20 on $2.00, Luna $0.02 on $0.20, GPT-5.3 Codex $0.175 on $1.75, ChatGPT (chat-latest) $0.50 on $5.00, GPT-5.6 Cyber $1.25 on $12.50 and the access-restricted GPT-Rosalind $0.50 on $5.00, verified September 14, 2026, and September 16 for Astra and Rosalind. Google, all ten Gemini models with a cache price, from Gemini 2.5 Flash-Lite at $0.01 on $0.10 to Gemini 2.5 Pro at $0.125 on $1.25, September 14, and Gemini Robotics ER 2 Preview at $0.10 on $1.00, September 16. Mistral’s two GLM rows, GLM 5.2 and GLM 5.3, both $0.14 on $1.40, September 14 and 16. And Anthropic for eleven of its thirteen models, where the pricing page says it in words: “A cache hit costs 10% of the standard input price”. Claude Opus 4.6 is $0.50 on $5.00, Claude Sonnet 5 $0.20 on $2.00, Claude Haiku 4.5 $0.10 on $1.00, September 14.
If you only remember one number, 0.1× is the right one. The rest of this post is about the eighteen rows where it is wrong.
The exceptions
Anthropic, 0.025× on two models. The same page carries a footnote: “Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. All other models use the standard 0.1x multiplier.” Claude Fable 5.1 and Claude Mythos 5.1 both cost $10.00 per million input tokens and $0.25 per million cached, a quarter of what the rule would give. Their predecessors, Claude Fable 5 and Claude Mythos 5, are $1.00 on the same $10.00, the ordinary 0.1×. Verified September 14 and 15, 2026.
DeepSeek, 0.033× and 0.02×. DeepSeek V4 Pro charges $0.044 per million cached tokens on $1.32 of input, and DeepSeek V4.1 Flash $0.006 on $0.30. That 0.02× is the lowest ratio among the list prices in the dataset; the only figure that matches it is Meta’s conditional price, below. Those are DeepSeek’s peak rates; outside Monday to Friday 01:00 to 04:00 and 06:00 to 10:00 UTC, every rate is half, cache included ($0.022 and $0.003). Verified September 14, 2026.
Meta, 0.12×, and a second price with a condition. Muse Spark 1.3 and Muse Spark 1.2 are $0.15 cached on $1.25 of input, read by hand in a browser on September 16, 2026 from the “Models and pricing” table on each model page. The same table lists a second SKU for each model, labelled “Used to improve our products”: $0.10 input, $0.002 cached and $0.20 output, which is 0.02× the input, but only if Meta may use your traffic to improve its products. This site publishes the row Meta labels “Not used to improve our products” as the list price, for the same reason it publishes OpenAI’s Standard tab and not Batch: a price that depends on giving something up is not the price. Until September 16 these two rows said “no published price”, because Meta’s pages show the table only to a browser and not to the HTTP client this site verifies with; that was a reading error here, not a change by Meta, and it is recorded on both model pages.
xAI, 0.15× to 0.25×. No single multiplier: Grok 4.6 is $0.50 cached on $2.00 of input, 0.25×; Grok 4.5 $0.30 on the same $2.00, 0.15×; Grok 4.3 and the three Grok 4.20 variants $0.20 on $1.25, 0.16×; Grok Build 0.1 $0.20 on $1.00, 0.20×. Verified September 14, 2026.
Alibaba, 0.107× to 0.2×, and two cache prices. In the Singapore region this site publishes, Qwen3.8 Max is $0.25 cached on $2.00 of input (0.125×), Qwen3.8 Flash $0.016 on $0.15 (0.107×), Qwen3.7 Plus $0.08 on $0.40 and Qwen3.7 Max $0.50 on $2.50 (0.2×), all verified September 11, 2026, and Qwen3.7 Flash $0.006 on $0.03 (0.2×), added and verified September 17. Those are the figures Alibaba labels “Input(Implicit Cache)”. The same pages list a second, lower figure for “Explicit Cache Read”, $0.17 for Qwen3.8 Max as read on September 15, 2026, which applies only to caches the caller creates explicitly and pays to create; the table stores the implicit one, the price paid without doing anything, and the explicit figure is quoted here from the page, not stored as a field.
Writing to the cache is not free
Four providers charge for putting content into the cache, in two shapes: a write price above the input price, at Anthropic and OpenAI, and a charge for keeping or creating the cache, by the hour at Google and per explicit creation at Alibaba.
Anthropic bills a cache write above the input price: 1.25× for a five-minute cache and 2× for a one-hour cache, on all thirteen models. For Claude Opus 4.6 that is $6.25 and $10.00 against $5.00 of input; for Fable 5.1, $12.50 and $20.00 against $10.00. OpenAI bills cache writes on the GPT-5.6 family at 1.25× the input price, as Sol’s model page states: “Cache writes are billed at 1.25x the uncached input token rate”. Sol $5.00, Terra $2.50, Luna $0.25. Verified September 14, 2026.
Google charges storage by the hour instead of a write fee. On Gemini 2.5 Pro and Gemini 3.1 Pro Preview it is $4.50 per million cached tokens per hour, verified September 14, 2026. The Flash models have a storage price too: as read on Google’s page on September 15, 2026, $1.00 per million tokens per hour, and $0.50 through December 31, 2026 on the three promotional ones. That figure is quoted from the page with its date; it is not a field in this dataset. Alibaba, for its explicit caches, charges an “Explicit Cache Creation” price, $2.50 per million tokens for Qwen3.8 Max as read on September 15, 2026, above the $2.00 input price; the same applies, a quote with a date, not a field.
The expensive request is the one that writes the cache, not the first hit after it. At 1.25×, the write costs a quarter more than a plain request for the same tokens, and a single cache read at 0.1× recovers it: write plus one read comes to 1.35× the input price, against 2× for two plain requests. At 2×, one read is not enough (2.1× against 2×) and it takes two (2.2× against 3×). At Anthropic’s 0.025× rate on Fable 5.1 and Mythos 5.1 the count is the same: one read recovers a 1.25× write, two recover a 2× write. With Google’s storage, an idle cache keeps costing money instead.
Three cache prices that are promotional
GPT-5.6 Sol’s $0.40 is part of a promotional price that OpenAI states is available at least through November 21, 2026, with no list figure printed anywhere; the cache price will move with the rest when the promotion ends, and this site does not guess to what. Gemini 3.6 Flash, 3.7 Flash and 3.8 Flash are $0.075 cached until December 31, 2026, and Google prints the price after: “$0.15 starting January 1, 2027”. So is Gemini Robotics ER 2 Preview: “$0.10 through December 31, 2026. $0.20 starting January 1, 2027”. All are marked as promotional on their pages and in the table.
What this site does not check every week
The weekly detector that verifies this dataset compares each row’s input and output price against the provider’s page. It does not read the cache column. Meta is outside that detector altogether: its pages serve the price table only to a browser, so the two Muse Spark rows are checked by hand only, input, output and cache alike, and no weekly run touches them. The verification date on a row is the date those two figures were confirmed. The 49 cache prices in this post, plus the write, storage and tier figures, were checked by hand, 42 of them on September 15, 2026, six on September 16 (Meta’s two and the four rows added to the dataset that day) and one on September 17 (Qwen3.7 Flash), each against the page the row cites, and one of them needed a second page: Claude Mythos 5.1’s row cites its model overview, which prints only input and output, so its cache prices come from Anthropic’s general pricing page, and the row now says so. Until the detector covers the column, treat the cache figures as verified on those dates, not on the row’s.
Using the number
The cost calculator takes the cache-hit rate you set as a share of the input tokens, charges that share at the model’s cached-input price and the rest at its input price; for a model that publishes no cache price, that share is charged at the input price too. When the average prompt size you enter passes a model’s long-context threshold, it uses that tier’s prices, cache price included. It adds no write fee, no storage, no off-peak rate and no regional price. For a workload that reuses a long system prompt, the ratio in this post is the number that moves the bill, and it is 0.1× in four providers, 0.025× in two Anthropic models, 0.02× to 0.033× at DeepSeek, 0.12× at Meta, and between 0.107× and 0.25× at xAI and Alibaba.
Frequently asked questions
How much does a prompt cache hit cost?
Usually 10% of the input price: 31 of the 49 current models with a cache-hit price sit at exactly 0.1×, at OpenAI, Google, Mistral's two GLM rows and eleven of Anthropic's thirteen models, as of September 17, 2026. The exceptions run from 0.02× at DeepSeek to 0.25× at xAI.
Which models have the cheapest cache hits?
Among list prices, DeepSeek V4.1 Flash at $0.006 per million cached tokens on $0.30 of input, 0.02×, followed by Claude Fable 5.1 and Claude Mythos 5.1 at 0.025×, $0.25 on $10.00. Both were verified on September 14 and 15, 2026.
Does writing to the prompt cache cost extra?
At Anthropic a cache write costs 1.25× the input price for a five-minute cache and 2× for a one-hour cache, and OpenAI bills writes on the GPT-5.6 family at 1.25×. Google charges storage by the hour instead, $4.50 per million cached tokens per hour on Gemini 2.5 Pro.
How many cache reads pay back a cache write?
One read at 0.1× recovers a 1.25× write: write plus one read comes to 1.35× the input price, against 2× for two plain requests. A 2× write needs two reads, 2.2× against 3×.
Written with AI assistance from this site's own verified pricing dataset and reviewed by Álvaro Lucero Rufino, who is responsible for what stays published here. Every figure links to the page where it was verified, with the date it was read; if a number here and a number on that page ever disagree, the page is the one to trust. Automated publishing was discontinued in September 2026; posts that no longer met that standard were removed rather than left up.