Data
LLM API pricing, compared
Every row links back to the provider's own pricing page. No estimates, no numbers pulled from memory - if we couldn't verify it, it says so instead of guessing.
Long prompts cost more on some models: the cheapest APIs for a 300,000-token prompt, tiers applied. Every recorded change: price changes. By release date: new models.
Data verified between September 11 and October 1, 2026 against each provider's official pricing page; 4 of 72 models still carry the oldest date. Each row shows its own. "Verified" also covers rows where we confirmed the provider publishes no per-token price. Methodology
Know your monthly volume? Calculate your actual cost across every model →
| Model | Provider | Input $/1M | Output $/1M | Cached input | Context | Status | Source |
|---|---|---|---|---|---|---|---|
| Qwen3.7-Flash | Alibaba (Qwen) | $0.03 | $0.13 | $0.006 | 1M | current | source |
| Command R7B | Cohere | $0.0375 | $0.15 | — | 128K | current | source |
| GPT-6 Luna | OpenAI | $0.10 | $0.50 | $0.01 | 1.1M | current | source |
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | $0.01 | 1.0M | current | source | |
| Ministral 3 3B | Mistral | $0.10 | $0.10 | — | 256K | current | source |
| Mistral Small 4 | Mistral | $0.15 | $0.60 | — | 256K | current | source |
| Ministral 3 8B | Mistral | $0.15 | $0.15 | — | 256K | current | source |
| Command R (08-2024) | Cohere | $0.15 | $0.60 | — | 128K | current | source |
| Qwen3.8-Flash | Alibaba (Qwen) | $0.15 | $0.47 | $0.016 | 1M | current | source |
| Qwen3.8-Omni-Flash | Alibaba (Qwen) | $0.15 | $0.47 | $0.016 | 1M | current | source |
| GPT-5.6 Luna | OpenAI | $0.20 | $1.20 | $0.02 | 1.1M | current | source |
| Ministral 3 14B | Mistral | $0.20 | $0.20 | — | 256K | current | source |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | $0.03 | 1.0M | current | source | |
| Gemini 2.5 Flash | $0.30 | $2.50 | $0.03 | 1.0M | current | source | |
| Codestral | Mistral | $0.30 | $0.90 | — | 128K | current | source |
| DeepSeek V4.1 Flash | DeepSeek | $0.30 | $1.20 | $0.006 | 1M | current | source |
| Qwen3.7-Plus | Alibaba (Qwen) | $0.40 | $1.60 | $0.08 | 1M | current | source |
| Mistral Large 3 | Mistral | $0.50 | $1.50 | — | 256K | current | source |
| Gemini 3.8 Flash promo until 31/12 | $0.75 | $3.75 | $0.075 | 1.0M | current | source | |
| Gemini 3.7 Flash promo until 31/12 | $0.75 | $3.75 | $0.075 | 1.0M | current | source | |
| Gemini 3.6 Flash promo until 31/12 | $0.75 | $3.75 | $0.075 | 1.0M | current | source | |
| Claude Haiku 4.5 | Anthropic | $1.00 | $5.00 | $0.10 | 200K | current | source |
| Gemini Robotics ER 2 Preview promo until 31/12 | $1.00 | $5.00 | $0.10 | 131K | current | source | |
| Gemini Robotics ER 2 Streaming Preview promo until 31/12 | $1.00 | $5.00 | — | 131K | current | source | |
| Grok Build 0.1 | xAI | $1.00 | $2.00 | $0.20 | 256K | current | source |
| Gemini 2.5 Pro | $1.25 | $10.00 | $0.125 | 1.0M | current | source | |
| Grok 4.3 | xAI | $1.25 | $2.50 | $0.20 | 1M | current | source |
| Grok 4.20 (reasoning) | xAI | $1.25 | $2.50 | $0.20 | 1M | current | source |
| Grok 4.20 (non-reasoning) | xAI | $1.25 | $2.50 | $0.20 | 1M | current | source |
| Grok 4.20 Multi-agent | xAI | $1.25 | $2.50 | $0.20 | 1M | current | source |
| Muse Spark 1.3 | Meta | $1.25 | $4.25 | $0.15 | 1M | current | source |
| Muse Spark 1.2 | Meta | $1.25 | $4.25 | $0.15 | 1M | current | source |
| DeepSeek V4 Pro | DeepSeek | $1.32 | $3.96 | $0.044 | 1M | current | source |
| GLM 5.2 | Mistral | $1.40 | $4.40 | $0.14 | 1M | current | source |
| GLM 5.3 | Mistral | $1.40 | $4.40 | $0.14 | 1M | current | source |
| Gemini 3.5 Flash | $1.50 | $9.00 | $0.15 | 1.0M | current | source | |
| Mistral Medium 3.5 | Mistral | $1.50 | $7.50 | — | 256K | current | source |
| GPT-5.3 Codex | OpenAI | $1.75 | $14.00 | $0.175 | 400K | current | source |
| GPT-6.1 Sol | OpenAI | $2.00 | $10.00 | $0.10 | 1.1M | current | source |
| GPT-6 Sol | OpenAI | $2.00 | $10.00 | $0.20 | 1.1M | current | source |
| GPT-5.6 Terra | OpenAI | $2.00 | $12.00 | $0.20 | 1.1M | current | source |
| Claude Sonnet 5.5 | Anthropic | $2.00 | $10.00 | $0.20 | 1M | current | source |
| Claude Sonnet 5 | Anthropic | $2.00 | $10.00 | $0.20 | 1M | current | source |
| Gemini 3.1 Pro Preview | $2.00 | $12.00 | $0.20 | 1.0M | current | source | |
| Grok 4.6 | xAI | $2.00 | $6.00 | $0.50 | 500K | current | source |
| Grok 4.7 | xAI | $2.00 | $6.00 | $0.50 | 500K | current | source |
| Grok 4.5 | xAI | $2.00 | $6.00 | $0.30 | 500K | current | source |
| Qwen3.8-Max | Alibaba (Qwen) | $2.00 | $6.00 | $0.25 | 1M | current | source |
| Command A | Cohere | $2.50 | $10.00 | — | 256K | current | source |
| Command R+ (08-2024) | Cohere | $2.50 | $10.00 | — | 128K | current | source |
| Qwen3.7-Max | Alibaba (Qwen) | $2.50 | $7.50 | $0.50 | 1M | current | source |
| Claude Sonnet 4.6 | Anthropic | $3.00 | $15.00 | $0.30 | 1M | current | source |
| GPT-5.6 Sol promo at least through 21/11 | OpenAI | $4.00 | $20.00 | $0.40 | 1.1M | current | source |
| Claude Opus 5.5 | Anthropic | $4.00 | $20.00 | $0.20 | 1M | current | source |
| GPT-Rosalind | OpenAI | $5.00 | $25.00 | $0.50 | — | current | source |
| ChatGPT (chat-latest) | OpenAI | $5.00 | $30.00 | $0.50 | 400K | current | source |
| Claude Opus 5 | Anthropic | $5.00 | $25.00 | $0.50 | 1M | current | source |
| Claude Opus 4.8 | Anthropic | $5.00 | $25.00 | $0.50 | 1M | current | source |
| Claude Opus 4.7 | Anthropic | $5.00 | $25.00 | $0.50 | 1M | current | source |
| Claude Opus 4.6 | Anthropic | $5.00 | $25.00 | $0.50 | 1M | current | source |
| Claude Opus 4.5 | Anthropic | $5.00 | $25.00 | $0.50 | 200K | current | source |
| GPT-6 Astra | OpenAI | $10.00 | $50.00 | $1.00 | 1.1M | current | source |
| Claude Fable 5.1 | Anthropic | $10.00 | $50.00 | $0.25 | 1M | current | source |
| Claude Mythos 5.1 | Anthropic | $10.00 | $50.00 | $0.25 | 1M | current | source |
| Claude Fable 5 | Anthropic | $10.00 | $50.00 | $1.00 | 1M | current | source |
| Claude Mythos 5 | Anthropic | $10.00 | $50.00 | $1.00 | 1M | current | source |
| GPT-5.6 Cyber | OpenAI | $12.50 | $75.00 | $1.25 | 400K | current | source |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | $0.025 | 1.0M | deprecated | source | |
| Command Light | Cohere | $0.30 | $0.60 | — | 4K | deprecated | source |
| Gemini 3 Flash Preview | $0.50 | — | $0.05 | — | deprecated | source | |
| Command R (03-2024) | Cohere | $0.50 | $1.50 | — | 128K | deprecated | source |
| Command | Cohere | $1.00 | $2.00 | — | 4K | deprecated | source |
| Claude Sonnet 4.5 | Anthropic | $3.00 | $15.00 | $0.30 | 200K | deprecated | source |
| Command R+ (04-2024) | Cohere | $3.00 | $15.00 | — | 128K | deprecated | source |
| Muse Spark 1.1 no published price | Meta | — | — | — | 1M | current | source |
| Command A+ no published price | Cohere | — | — | — | 128K | current | source |
| Command A Reasoning no published price | Cohere | — | — | — | 256K | current | source |
| Command A Vision no published price | Cohere | — | — | — | 128K | current | source |
| Command A Translate no published price | Cohere | — | — | — | 8K | current | source |
5 current models above show "no data" instead of a price - we couldn't verify a USD per-token number for them, and we don't guess:
- Cohere (4): Cohere publishes per-token prices for Command A, Command R7B and the Command R family, but not for Command A+, Command A Reasoning, Command A Vision or Command A Translate. Their docs pages say each is "free until rate limits are reached" on trial and production keys, and that production use goes through Cohere's Model Vault or through sales, with custom pricing. Checked September 12, 2026; see the Command A+ docs .
- Meta Muse Spark 1.1 (1): Meta's page for this model shows a benchmark table and no per-token price, and Meta's pricing tables for the later Muse Spark versions do not list it. We do not take figures from third parties. Checked by hand in a browser on September 16, 2026; see the model page .
Alibaba (Qwen) prices by region. The 6 Alibaba (Qwen) rows above show the Singapore (International) region: the one Alibaba labels international and the one every reseller in the provider comparison mirrors. Beijing and Alibaba's Global regions (Frankfurt, Virginia, Tokyo, Hong Kong) charge less:
- Qwen3.7-Flash: $0.028 / $0.11 instead of $0.03 / $0.13
- Qwen3.8-Flash: $0.113 / $0.382 instead of $0.15 / $0.47
- Qwen3.8-Omni-Flash: $0.113 / $0.382 instead of $0.15 / $0.47
- Qwen3.7-Plus: $0.276 / $1.101 instead of $0.40 / $1.60
- Qwen3.8-Max: $1.65 / $4.951 instead of $2.00 / $6.00
- Qwen3.7-Max: $1.65 / $4.951 instead of $2.50 / $7.50
Methodology
The table covers models charged per token with text output and a price published by the maker. Audio, image, video and embedding models, and anything priced per minute, are left out even when the same page prices them. For Alibaba, whose catalog keeps many older Qwen generations on sale, the table covers the current generations, Qwen3.7 and Qwen3.8. Every price above was read directly from the provider's own pricing documentation on the verification date shown - not estimated, not carried over from a previous check, not sourced from a third-party aggregator. "Verified" means one of two things, and the row says which: the price was read on the provider's official page, or we confirmed that the provider publishes no per-token price at all, in which case the row shows "no published price" with the date that absence was checked. A row we could not check either way says "no data". No row ever shows a guessed or third-party number. Where a provider prices by region, the row names the region it shows and the other level is published next to it. List prices are taken without promotions: Alibaba's pages, for example, state that "this page only shows the original pricing for model API calls, excluding any limited-time promotions", and that is the figure we record.
The current rows are also available as files, in JSON and CSV, together with the record of price changes, under a CC BY 4.0 licence: see the dataset page for the fields and how to cite them.
Analysis
What the dataset shows when it is read across providers.
- How LLM API pricing worksThe guide: tokens, cache reads and writes, long-prompt tiers, Batch and Fast mode, region, peak hours, promotions and hosts, each on a real row of the dataset.
- Same List Price, Different Bill: Five Pairs of LLM APIs Where the Workload Decides What You PayClaude Sonnet 5.5 and GPT-6.1 Sol both list at $2 / $10, yet a long-document workload costs $250 on one and $475 on the other. Five pairs, every assumption shown.
- Prompt Caching Is Not 10% Everywhere: What a Cache Hit Costs Across Eight Providers49 of the 64 current priced models publish a cache-hit price; four providers charge 10% of input. Anthropic, DeepSeek, xAI, Alibaba and Meta differ, with dates.
- The Second Price List: What Batch, Flex, Priority and Data-Sharing Discounts Cost Across Seven ProvidersHalf price to wait, double to skip the queue, 12.5× less input to share data: the second price lists of seven LLM providers, read by hand on September 16, 2026.
- Five Things Decide What You Pay for a Model, and Only One of Them Is the ModelHost, time of day, promotions, region and request size move the price of the same model by anything from 1.2× to 10×. Five verified cases, with dates and sources.
- OpenAI's Pricing Page Has Four Tabs That Look the Same in Plain Text. We Published the Wrong One for Seven Days.From September 3 to 11, 2026 this site listed GPT-5.6 Sol, Terra and Luna at half their price, read from the Batch tab. What went wrong, what caught it, and what is still not caught.