Same List Price, Different Bill: Five Pairs of LLM APIs Where the Workload Decides What You Pay
Claude Sonnet 5.5 and GPT-6.1 Sol cost the same on paper: $2.00 per million input tokens and $10.00 per million output tokens. Send each of them 100 million tokens a month in prompts of about 300,000 tokens, and the bill is $250.00 on one and $475.00 on the other. Nothing in the list price says which. The difference sits in the lines under it: the price of a cached token, and a higher price that starts above a prompt length.
The pricing table has 69 current models with a verified price. Among them, 21 pairs from different makers have exactly the same input and output price, and 2 more are within 10% of each other on both. This post takes five of those pairs and prices each one under the same set of workloads, to show which line of the price list decides the bill. It compares price only; nothing here says anything about what the models do.
How every figure below is calculated
Each cost is a month of usage priced with the formula of the cost calculator, on the prices of each model’s row in this site’s dataset:
- Input tokens are split into cache hits, charged at the model’s cached-input price, and the rest, charged at its input price. Where a model publishes no cache price, cached tokens are charged as normal input.
- Output tokens are charged at the output price.
- If the average prompt is longer than a model’s long-prompt threshold, that tier’s prices apply to the whole volume. That is the calculator’s rule, and the one the long-prompt ranking uses.
- Cache writes, Batch prices and off-peak rates are not in the formula. Where they change the comparison, they are given separately.
- The long-document workloads fit every model compared here. At 100M input tokens in prompts of 300,000, a month is about 333 requests, and 5M output tokens make about 15,000 tokens of answer each: some 315,000 tokens per request. The smallest context window among the ten models in this dataset is Grok 4.7’s, 500,000 tokens; the others have 1,000,000 or more.
The six workloads, the same for every pair:
| Workload | Input tokens a month | Output tokens a month | Share of input from cache | Average prompt |
|---|---|---|---|---|
| Short prompts | 10M | 2M | 0% | 2,000 tokens |
| Cached system prompt | 100M | 10M | 50% | 20,000 tokens |
| Heavy cache | 100M | 10M | 90% | 20,000 tokens |
| Documents of 250K | 100M | 5M | 0% | 250,000 tokens |
| Documents of 300K | 100M | 5M | 0% | 300,000 tokens |
| Documents of 300K, cached | 100M | 5M | 50% | 300,000 tokens |
Each one opens in the calculator with its values filled in: short, cached, heavy cache, 250K, 300K, 300K cached.
Claude Sonnet 5.5 and GPT-6.1 Sol: the cache and the threshold
Claude Sonnet 5.5 (verified September 29, 2026) and GPT-6.1 Sol (verified September 30, 2026) share the list price, $2.00 / $10.00. They differ in two lines. A cache hit costs $0.20 on Sonnet 5.5, 10% of its input price, and $0.10 on GPT-6.1 Sol, 5%. And GPT-6.1 Sol has a long-prompt tier: above 272,000 tokens of prompt it charges $4.00 input, $15.00 output and $0.20 for a cache hit. This dataset records no long-prompt tier for Sonnet 5.5.
| Workload | Claude Sonnet 5.5 | GPT-6.1 Sol |
|---|---|---|
| Short prompts | $40.00 | $40.00 |
| Cached system prompt | $210.00 | $205.00 |
| Heavy cache | $138.00 | $129.00 |
| Documents of 250K | $250.00 | $250.00 |
| Documents of 300K | $250.00 | $475.00 (tier) |
| Documents of 300K, cached | $160.00 | $285.00 (tier) |
With short prompts and no cache, the two bills are identical. Once half the input comes from the cache, GPT-6.1 Sol is $5.00 cheaper, 2.4%; at 90% it is $9.00 cheaper, 6.5%. The gap grows with the cache share, because the cached token is the only price where the two differ below the threshold.
Above 272,000 tokens the order flips and the gap is much wider: $475.00 against $250.00, 1.9 times, because GPT-6.1 Sol’s tier doubles the input price and raises output by half. With half the input cached, GPT-6.1 Sol’s tier cache price ($0.20) equals Sonnet 5.5’s, and the gap stays: $285.00 against $160.00.
The rule that falls out of the two rows is simple. At prompts up to 272,000 tokens, each of the three prices the formula uses (input, cache hit, output) is equal or lower on GPT-6.1 Sol, so under this formula it never costs more. Above that length, each of the three is equal or lower on Sonnet 5.5, so Sonnet 5.5 never costs more. The workload’s prompt length decides which one is cheaper; the cache share decides by how much.
Two lines the formula leaves out do not change this. Writing to the cache costs $2.50 per million tokens on both below the threshold (Sonnet 5.5’s five-minute cache; its one-hour cache costs $4.00), and $5.00 on GPT-6.1 Sol above it. On the Batch price list both charge half, $1.00 / $5.00, so the short workload costs $20.00 on each. OpenAI’s Batch list also prints a cached price for GPT-6.1 Sol, $0.05; the Anthropic figures this site records for Sonnet 5.5’s Batch list have no cached price, so a cached Batch workload is not compared here.
GPT-6 Astra and Claude Fable 5.1: one never costs more
GPT-6 Astra and Claude Fable 5.1 both list at $10.00 / $50.00 (both verified September 28, 2026). Here the cache runs the other way: a hit costs $1.00 on Astra, 10% of input, and $0.25 on Fable 5.1, 2.5%. Astra also has a tier above 272,000 tokens, $20.00 / $75.00, and Fable 5.1 has none recorded.
| Workload | GPT-6 Astra | Claude Fable 5.1 |
|---|---|---|
| Short prompts | $200.00 | $200.00 |
| Cached system prompt | $1,050.00 | $1,012.50 |
| Heavy cache | $690.00 | $622.50 |
| Documents of 250K | $1,250.00 | $1,250.00 |
| Documents of 300K | $2,375.00 (tier) | $1,250.00 |
| Documents of 300K, cached | $1,475.00 (tier) | $762.50 |
Each of the three prices the formula uses is equal or lower on Fable 5.1, so no workload in this formula makes Astra cheaper. They tie only when nothing is cached and prompts stay under 272,000 tokens. Their Batch lists are the same, $5.00 / $25.00. Writing to the cache is where Fable 5.1 is not lower: $12.50 for its five-minute cache, the same as Astra’s cache write, and $20.00 for its one-hour cache.
Grok 4.7 and Qwen3.8-Max: cache, threshold and region
Grok 4.7 (verified September 28, 2026) and Qwen3.8-Max (verified September 11, 2026) both charge $2.00 / $6.00. A cache hit is $0.50 on Grok 4.7 and $0.25 on Qwen3.8-Max. Grok 4.7 has a tier from 200,000 tokens of prompt, $4.00 / $12.00, and a 500,000-token context window; Qwen3.8-Max has no tier recorded.
There is a third line: Qwen3.8-Max’s price is Alibaba’s Singapore (International) price, the one this site publishes. In Beijing and Alibaba’s Global regions it costs $1.65 / $4.951, with cache hits at $0.206.
| Workload | Grok 4.7 | Qwen3.8-Max, Singapore | Qwen3.8-Max, Beijing and Global |
|---|---|---|---|
| Short prompts | $32.00 | $32.00 | $26.40 |
| Cached system prompt | $185.00 | $172.50 | $142.31 |
| Heavy cache | $125.00 | $102.50 | $84.55 |
| Documents of 250K | $460.00 (tier) | $230.00 | $189.76 |
| Documents of 300K | $460.00 (tier) | $230.00 | $189.76 |
| Documents of 300K, cached | $310.00 (tier) | $142.50 | $117.56 |
Grok 4.7’s threshold is lower than GPT-6.1 Sol’s, so its tier already applies to the 250,000-token documents: $460.00 against $230.00, twice as much. As with the pair above, one model is never the more expensive of the two: at the Singapore price, Qwen3.8-Max ties Grok 4.7 only with short, uncached prompts.
GPT-5.6 Terra and Gemini 3.1 Pro Preview: only the threshold differs
GPT-5.6 Terra and Gemini 3.1 Pro Preview (both verified September 28, 2026) are the closest pair in the table. Both charge $2.00 / $12.00 with cache hits at $0.20, and both have a tier at $4.00 / $18.00 with cache hits at $0.40. The only difference is where the tier starts: above 272,000 tokens for GPT-5.6 Terra, above 200,000 for Gemini 3.1 Pro Preview.
| Workload | GPT-5.6 Terra | Gemini 3.1 Pro Preview |
|---|---|---|
| Short prompts | $44.00 | $44.00 |
| Cached system prompt | $230.00 | $230.00 |
| Heavy cache | $158.00 | $158.00 |
| Documents of 250K | $260.00 | $490.00 (tier) |
| Documents of 300K | $490.00 (tier) | $490.00 (tier) |
| Documents of 300K, cached | $310.00 (tier) | $310.00 (tier) |
Five of the six workloads cost the same to the cent. The one that does not, prompts of 250,000 tokens, costs 1.88 times as much on Gemini 3.1 Pro Preview, because it falls between the two thresholds. For this pair the whole comparison is one question: are the prompts between 200,000 and 272,000 tokens long? One line is missing from the comparison: GPT-5.6 Terra prints a cache-write price, $2.50, and Gemini 3.1 Pro Preview’s row in this dataset records none, so cache writes are not compared.
Muse Spark 1.3 and DeepSeek V4 Pro: nearly the same list, and the hour
Muse Spark 1.3 (verified September 16, 2026) lists at $1.25 / $4.25 and DeepSeek V4 Pro (verified September 28, 2026) at $1.32 / $3.96: input 5.6% apart, output 7.3% apart, in opposite directions. Cache hits cost $0.15 on Muse Spark 1.3 and $0.044 on DeepSeek V4 Pro. Neither has a long-prompt tier recorded.
DeepSeek V4 Pro’s list price is its peak rate, which applies Monday to Friday from 01:00 to 04:00 and from 06:00 to 10:00 UTC, 35 of the 168 hours in a week. The rest of the time it charges half: $0.66 / $1.98, with cache hits at $0.022 (tariff read September 14, 2026).
| Workload | Muse Spark 1.3 | DeepSeek V4 Pro, list (peak) | DeepSeek V4 Pro, off-peak |
|---|---|---|---|
| Short prompts | $21.00 | $21.12 | $10.56 |
| Cached system prompt | $112.50 | $107.80 | $53.90 |
| Heavy cache | $68.50 | $56.76 | $28.38 |
| Documents of 300K | $146.25 | $151.80 | $75.90 |
| Documents of 300K, cached | $91.25 | $88.00 | $44.00 |
(The 250,000-token documents cost the same as the 300,000-token ones here: neither model has a tier.)
At list prices this pair does not have a single winner. Muse Spark 1.3 is cheaper in both workloads where nothing is cached, because its input price is lower; DeepSeek V4 Pro is cheaper as soon as the cache carries part of the input, because its cache hit costs less than a third of Muse Spark 1.3’s. At off-peak hours DeepSeek V4 Pro costs half its list price in every workload.
The five pairs side by side
Monthly cost under the workloads above, at list price. “Tier” marks a long-prompt tier applied to the whole volume.
| Pair (list price, input / output) | Short prompts | Cached 50% | Cached 90% | Documents of 250K | Documents of 300K |
|---|---|---|---|---|---|
| Claude Sonnet 5.5 / GPT-6.1 Sol ($2 / $10 both) | $40.00 / $40.00 | $210.00 / $205.00 | $138.00 / $129.00 | $250.00 / $250.00 | $250.00 / $475.00 tier |
| GPT-6 Astra / Claude Fable 5.1 ($10 / $50 both) | $200.00 / $200.00 | $1,050.00 / $1,012.50 | $690.00 / $622.50 | $1,250.00 / $1,250.00 | $2,375.00 tier / $1,250.00 |
| Grok 4.7 / Qwen3.8-Max ($2 / $6 both) | $32.00 / $32.00 | $185.00 / $172.50 | $125.00 / $102.50 | $460.00 tier / $230.00 | $460.00 tier / $230.00 |
| GPT-5.6 Terra / Gemini 3.1 Pro Preview ($2 / $12 both) | $44.00 / $44.00 | $230.00 / $230.00 | $158.00 / $158.00 | $260.00 / $490.00 tier | $490.00 tier / $490.00 tier |
| Muse Spark 1.3 / DeepSeek V4 Pro ($1.25 / $4.25, $1.32 / $3.96) | $21.00 / $21.12 | $112.50 / $107.80 | $68.50 / $56.76 | $146.25 / $151.80 | $146.25 / $151.80 |
What the five pairs have in common
In every pair, short prompts with nothing cached cost the same or within a few cents: that is the workload the list price describes. Each of the other lines moves the bill in its own direction. A cheaper cache hit widens the gap in proportion to the share of input it covers, which is why the difference between Sonnet 5.5 and GPT-6.1 Sol goes from 2.4% at half cached to 6.5% at 90%. A long-prompt tier does not move the bill gradually: below the threshold it does nothing, above it the whole workload is priced higher, which is why the largest gaps in the summary table, up to 2.0 times, all come from a prompt length crossing a threshold. And two lines only one of the ten models has, Alibaba’s regions and DeepSeek’s off-peak hours, lower the bill of that model alone.
To price a workload of your own, the cost calculator applies the same formula to every model in the table. The ranking for a 300,000-token prompt applies every tier to one long request, and How LLM API pricing works explains each line of a price list: tokens, cache, tiers, Batch, region, peak hours, promotions and hosts.
Frequently asked questions
Do LLM APIs with the same list price cost the same?
Only for short prompts with nothing cached. Claude Sonnet 5.5 and GPT-6.1 Sol both list at $2.00 / $10.00, yet 100 million input tokens a month in prompts of about 300,000 tokens cost $250.00 on Sonnet 5.5 and $475.00 on GPT-6.1 Sol.
Is Claude Sonnet 5.5 or GPT-6.1 Sol cheaper?
With prompts up to 272,000 tokens GPT-6.1 Sol never costs more, and it is cheaper once part of the input is cached ($0.10 per cache hit against $0.20). Above 272,000 tokens its long-prompt tier of $4.00 / $15.00 applies and Sonnet 5.5 never costs more.
Is GPT-6 Astra or Claude Fable 5.1 cheaper?
Claude Fable 5.1 never costs more under the calculator's formula: both list at $10.00 / $50.00, but its cache hit costs $0.25 against $1.00, and it has no long-prompt tier recorded while Astra charges $20.00 / $75.00 above 272,000 tokens. They tie only when nothing is cached and prompts stay under 272,000 tokens.
What is the difference between GPT-5.6 Terra and Gemini 3.1 Pro Preview pricing?
Only where the long-prompt tier starts: both charge $2.00 / $12.00 with cache hits at $0.20 and a tier at $4.00 / $18.00, but it begins above 272,000 tokens on GPT-5.6 Terra and above 200,000 on Gemini 3.1 Pro Preview. Prompts of 250,000 tokens therefore cost 1.88 times as much on Gemini 3.1 Pro Preview.
Written with AI assistance from this site's own verified pricing dataset and reviewed by Álvaro Lucero Rufino, who is responsible for what stays published here. Every figure links to the page where it was verified, with the date it was read; if a number here and a number on that page ever disagree, the page is the one to trust. Automated publishing was discontinued in September 2026; posts that no longer met that standard were removed rather than left up.