LLM API pricing / Guide

How LLM API pricing works

· Figures from the dataset, prices verified up to October 1, 2026

An LLM API charges per million tokens, with one price for the tokens you send and a higher one, a median of 5× across the 67 current models with a price on this site, for the tokens the model writes. What a request finally costs then depends on six more things: whether part of the prompt is read from a cache, how long the prompt is, whether you can wait for the answer, which region and hour you call in, whether a promotion is running, and who serves the model. Each section below shows one of them on a real row of the pricing table, with the date it was verified and a link to the page where it can be checked.

Tokens: why input and output have two prices

Every price in this dataset is per million tokens (MTok), the units in which a model counts the text it reads and the text it writes. The prompt, the system instructions and any documents sent with them are input; the answer is output. The two are priced apart, and on 64 of the 67 priced models output costs more. This site records the two prices, not the reasons for the gap, and the gap varies: 26 models charge exactly 5× the input price for output, the median is 5×, and the range runs from 1× (Ministral 3 3B, Ministral 3 8B and Ministral 3 14B) to 8.33× (Gemini 3.5 Flash-Lite and Gemini 2.5 Flash).

Example. GPT-6.1 Sol costs $2.00 per million input tokens and $10.00 per million output tokens (verified September 30, 2026). A request with a 10,000-token prompt and a 1,000-token answer costs $0.02 for the prompt and $0.01 for the answer, $0.03 in all: the answer is 9.1% of the tokens and 33.3% of the cost. At the ends of the range, Ministral 3 14B charges $0.20 for both (verified September 28, 2026) and Gemini 2.5 Flash $0.30 in and $2.50 out (verified October 1, 2026). The cost calculator does this sum for a monthly volume on every model at once.

Prompt caching: reads and writes

When part of a prompt has been sent before, a maker that offers prompt caching charges a lower price for the tokens read from the cache; 8 of the 9 makers in the table publish one. 55 of the 67 priced models publish that cache-hit price, and it is not the same share of the input price everywhere. Google charges 10% on the 10 of its 11 models that list one, Meta charges 12% and Mistral charges 10% on the 2 of its 9 models that list one. Others vary by model: Alibaba (Qwen) 20% (3 models), 10.7% (2 models) and 12.5%, Anthropic 10% (11 models), 2.5% (2 models) and 5%, DeepSeek 2% and 3.3%, OpenAI 10% (10 models) and 5% and xAI 16% (4 models), 25% (2 models), 15% and 20%. Cohere lists no cache price for its current models.

Writing to the cache can cost extra. OpenAI and Anthropic print a cache-write price, 1.25× the input price on all 21 of their models that have one; 14 Anthropic models add a one-hour cache at 2× the input price.

Example. Claude Sonnet 5.5: input $2.00, cache read $0.20 (10%), five-minute cache write $2.50 and one-hour write $4.00, verified September 29, 2026. On Grok 4.7 a cache hit costs $0.50 against $2.00 of input, 25% (verified September 28, 2026). The post Prompt caching is not 10% everywhere goes through each maker's exceptions.

Long prompts: tiers by prompt length

19 of the 67 priced models charge more per token once the prompt passes a length the maker sets; the other 48 have no such tier recorded. The thresholds in the dataset are 32,000, 200,000, 256,000 and 272,000 tokens. Alibaba (Qwen) has a tier on 2 of its 6 models, Google has a tier on 2 of its 11 models, 2× input and 1.5× output above 200,000 tokens, OpenAI has a tier on 7 of its 11 models, 2× input and 1.5× output above 272,000 tokens and xAI has a tier on 8 of its 8 models, 2× input and 2× output above 200,000 tokens. Qwen3.7-Flash has more than one step.

Example. GPT-6.1 Sol costs $2.00 input and $10.00 output up to 272,000 tokens of prompt and $4.00 input and $15.00 output above, verified September 30, 2026. A 300,000-token prompt with a 2,000-token answer costs $1.23 at the tier price, against the $0.62 the base price would give. The long-prompt ranking applies every model's tier to that same request.

Batch, Flex and Fast mode

Next to the standard price, some makers publish second price lists that trade speed for money. This site records them only where the maker's page prints a figure or names the model: OpenAI Batch (half of Standard on every column; wait: results within 24 hours), OpenAI Flex (half of Standard on every column, the same figures as Batch; accept slower responses and occasional unavailability), OpenAI Fast mode (double Standard on every column; pay more to be scheduled first), OpenAI Ultrafast (the figures printed on the pricing page's Ultrafast tab; pay more for the fastest output, GPT-6 Astra only), Anthropic Batch (50% discount on both input and output tokens; wait: results when the batch completes or after 24 hours), Anthropic Fast mode (double Standard; pay more for faster output, research preview, first-party API only), Google Batch (50% of the standard interactive API cost; wait: designed to complete within 24 hours), Google Flex (50% cost reduction compared to standard rates; accept variable latency and best-effort availability), Google Priority (no multiplier printed; the page prints the prices; pay more to be prioritized above standard and Flex traffic, in Preview) and xAI Batch (20 % off all token types; wait: completion is best effort, most within 24 hours). They are not the list price, and the table never shows them in its price columns.

Example. GPT-6 Astra is $10.00 input and $50.00 output at the standard rate (verified September 28, 2026). OpenAI's pages print, for the same model: Batch $5.00 input and $25.00 output (read September 16, 2026); Flex $5.00 input and $25.00 output (read September 16, 2026); Fast mode $20.00 input and $100.00 output (read September 16, 2026); Ultrafast $60.00 input and $300.00 output (read September 30, 2026). From 0.5× to 6× the standard price for the same tokens. The post The second price list compares these lists across providers.

Region and time of day

6 priced models, all from Alibaba (Qwen), have a different price by region. This site shows the Singapore (International) price, the one Alibaba calls international, and publishes the other level next to it; the Singapore price is 1.07× to 1.52× the Beijing and Global one on input. Example: Qwen3.8-Flash costs $0.15 input and $0.47 output in Singapore (International), and $0.113 input and $0.382 output in Beijing and Global (verified September 11, 2026).

2 priced models, from DeepSeek, change price with the hour. The maker's list price applies Monday to Friday 01:00-04:00 and 06:00-10:00 UTC, 35 of the 168 hours in a week (21%); the rest of the time it charges 50% of it. Example: DeepSeek V4.1 Flash is $0.30 input and $1.20 output at peak and $0.15 input and $0.60 output off-peak (tariff read September 14, 2026). The table shows the list price. Both factors, with the cases that brought them to light, are in Five things decide what you pay for a model.

Promotions with an end date

6 priced models carry a promotional price today. The table shows the price you pay now, marked with the day the promotion ends, and, for 5 of them, the list price it returns to. For GPT-5.6 Sol the maker publishes no price for after the promotion, and none is shown. End dates: December 31, 2026 (5 models) and November 21, 2026 (1 model).

Example. Gemini Robotics ER 2 Preview costs $1.00 input and $5.00 output until December 31, 2026, and $2.00 input and $10.00 output after (verified October 1, 2026): the list price is 2× the promotional one. When a promotion ends, the change is recorded on the price changes page.

The maker's price and the hosts' price

Open-weights models are sold by more than one host. Of the 8 open-weights models this site prices on two or more hosts, 6 cost the same everywhere, within 1.15× on input, and 2 do not (Llama 3.3 70B Instruct, 10.4× and DeepSeek V4 Flash 0731, 2.33×). 4 of the 8 have a price from their own maker, and all 4 are in the same-price group.

Example. Kimi K3: DeepInfra $2.85 / $14.25 (September 30, 2026); Together AI $3.00 / $15.00 (October 1, 2026); Fireworks AI $3.00 / $15.00 (October 1, 2026); Moonshot AI (the maker, through OpenRouter) $3.00 / $15.00 (October 1, 2026). The widest gap is 1.05×. Inference provider pricing lists every model and host, When to compare hosts explains which ones are worth comparing, and OpenRouter pricing explains why the number in OpenRouter's catalog is not what a request pays.

Questions

How is LLM API usage priced?

Per million tokens, with one price for the tokens you send (input) and another for the tokens the model writes (output). Across the 67 current models with a price on this site, output costs a median 5× the input price; 26 of them charge exactly 5×, and the range runs from 1× to 8.33×.

How much does a cached input token cost?

55 of the 67 priced models publish a cache-hit price. The most common is 10% of the input price (33 models); the full range runs from 2% to 25%. OpenAI and Anthropic also charge for writing to the cache, at 1.25× the input price.

Do longer prompts cost more per token?

On 19 of the 67 priced models, yes: once the prompt passes a length set by the maker (32,000, 200,000, 256,000 and 272,000 tokens, depending on the model), a higher price per token applies. The step is set by the maker, from 2× to 3.33× the input price at the first threshold.

How much cheaper is the Batch API?

OpenAI, Anthropic and Google price Batch at half their standard rate, in exchange for asynchronous results within a 24-hour window. xAI's page says "20 % off all token types". Paying for speed goes the other way: OpenAI's and Anthropic's Fast mode cost double.

Does the same open-weights model cost the same on every host?

Mostly. Of the 8 open-weights models this site prices on two or more hosts, 6 cost the same everywhere (within 1.15× on input) and 2 do not: Llama 3.3 70B Instruct (10.4×) and DeepSeek V4 Flash 0731 (2.33×).

Where each figure lives on this site