pricingbatch-apillm-apimethodology

The Second Price List: What Batch, Flex, Priority and Data-Sharing Discounts Cost Across Seven Providers

The pricing table on this site publishes one price per model: Standard processing, immediate, on the paid tier, and, where the provider says so on its page (Meta and Google do), with your traffic not used to improve its products. Every provider also sells the same model at other prices, and each of those comes with a condition: wait, pay more, or let the provider use your traffic. This post is a reading of those second price lists, by hand, on September 16, 2026, from each provider’s pricing page and the documentation page of each modality; the quotes are the providers’ own words on that date. None of these figures is in the table or the calculator, which use only the Standard price, and the weekly check that verifies the table does not watch these lists.

Waiting: Batch and Flex

Five of the seven providers sell a slower lane at a lower price.

OpenAI prints four tabs, Standard, Batch, Flex and Fast mode, and Batch and Flex are half of Standard on every column. On GPT-6 Astra, $10.00 in and $50.00 out become $5.00 and $25.00; cached input drops from $1.00 to $0.50 and the cache write from $12.50 to $6.25. The Batch condition is time: “a clear 24-hour turnaround time”, and “for now, the completion window can only be set to 24h”; a batch that does not finish “eventually move[s] to an expired state; unfinished requests within that batch are cancelled”. The Flex condition is capacity: “lower costs for Responses or Chat Completions requests in exchange for slower response times and occasional resource unavailability”. One caveat this site has already written about: GPT-5.6 Sol’s Standard price of $4.00 / $20.00 is itself promotional, so its Batch and Flex figure of $2.00 / $10.00 is half of a promotion, not half of a list price. How those four tabs look in plain text, and what went wrong when this site read them, is in the post on OpenAI’s four tabs.

Anthropic states the rule in one sentence: “The Batch API allows asynchronous processing of large volumes of requests with a 50% discount on both input and output tokens.” The Batch table prices Claude Opus 5 at $2.50 / $12.50 against $5.00 / $25.00 and Claude Fable 5.1 at $5.00 / $25.00 against $10.00 / $50.00, and it covers every current Anthropic model in this dataset, thirteen rows, each of which now shows the discount on its page with the date it was checked. The condition is softer than OpenAI’s: “most batches finishing in less than 1 hour”, results available “when all messages have completed or after 24 hours, whichever comes first”, and “batches expire if processing does not complete within 24 hours”.

Google sells both lanes. Batch “is priced at 50% of the standard interactive API cost for the equivalent model”, with a “24-hour turnaround time” as the service level objective; Flex “offers a 50% cost reduction compared to standard rates, in exchange for variable latency and best-effort availability”, and “if there is a spike in standard traffic, Flex requests may be preempted or evicted”. On Gemini 2.5 Pro both lanes are $0.625 in and $5.00 out against $1.25 and $10.00; on Gemini 3.6 Flash, $0.375 and $1.875 against a Standard price that is itself promotional until December 31, 2026.

xAI publishes a discount that depends on the model: “20 %” for Grok 4.3 and the three Grok 4.20 variants, and “models not listed above have no batch discount”. It applies “to all token types”, and completion “is best effort and not guaranteed”, with most requests done “within 24 hours”. The resulting per-model prices sit behind a toggle on each model’s page that this reading could not open, so no batch figure is quoted here.

Mistral says “Batch High-volume processing, at half price, for maximum efficiency” on its pricing page and “a 50% discount” in its documentation, “ideal for tasks with high throughput requirements but low latency sensitivity or priority”. The per-model batch column is rendered by a script and could not be read, so, as for xAI, only the multiplier is quoted.

Paying to skip the queue: Priority and Fast

The same lever runs the other way. OpenAI charges double for what it now calls Fast mode: “Priority processing was renamed Fast mode on July 30, 2026”, and the tab prints Astra at $20.00 / $100.00 and Sol at $8.00 / $40.00. The premium buys scheduling, not a guarantee: “if your traffic ramps too fast, the system may downgrade some Fast mode requests to standard speeds and charge standard rates”.

Anthropic’s Fast mode is “in research preview”, on two models only, “Claude Opus 5 and Claude Opus 4.8 at premium pricing”: $10 / $50 against $5 / $25, on the first-party API only, and “not available with the Batch API”. A request to an unsupported model with the fast flag “run[s] at standard speed and [is] billed at standard rates”.

Google’s Priority tier is “in Preview” and “prioritized above standard API and Flex tier traffic”. The page prints prices and no multiplier: Gemini 2.5 Pro at $2.25 / $18.00 against $1.25 / $10.00, and Gemini 3.6 Flash at $1.35 / $6.75 against $0.75 / $3.75. Both work out to 1.8× Standard; that ratio is this site’s arithmetic, not a figure Google publishes, and for 3.6 Flash it is computed over a Standard price that is itself promotional until December 31, 2026.

xAI prints the multiplier: “Priority requests are billed at a 2x premium over standard rates”, on “all token types”, with one honest detail: “You are only billed at the priority rate when the response confirms “service_tier”: “priority”. If the request is served at the default tier instead, standard rates apply.” Mistral lists a Priority tier on its pricing page and publishes no figure for it.

Giving up data: Meta’s contributor price and Google’s free tier

The cheapest lane of all is paid in traffic. Meta lists two SKUs for Muse Spark 1.3 and 1.2: the one labelled “Not used to improve our products” at $1.25 in, $0.15 cached and $4.25 out, and a “contributor” SKU labelled “Used to improve our products” at $0.10, $0.002 and $0.20, 12.5× less on input, read by hand in a browser on September 16, 2026. This site publishes the first as the list price, for the reason given in the post on cache prices: a price that depends on giving something up is not the price.

Google draws the same line inside its pricing page. Every model block has a Free Tier and a Paid Tier column, and the last row of each block, “Used to improve our products”, reads “Yes” under Free Tier and “No” under Paid Tier. On Gemini 2.5 Pro the free column says “Free of charge” for Standard input and output, and “Not available” for Batch and Flex; the paid column is the $1.25 / $10.00 in the table.

Another hour: DeepSeek

DeepSeek’s second list is a clock rather than a queue: every rate is half outside its peak window, “01:00 - 04:00 and 06:00 - 10:00 UTC, Monday through Friday”, so DeepSeek V4 Pro costs $0.66 / $1.98 for 133 of the 168 hours of the week against a list price of $1.32 / $3.96. The figures, the probes that showed which hosts follow the clock and what it means for the table are in the post on the five things that decide a price, section 2, and are not repeated here.

What this site publishes, and what it does not

The table and the calculator use Standard prices only: immediate, paid tier, no data sharing, peak rate where a tariff exists. Nothing above is a list price, and none of it is a field in the dataset except one: Anthropic’s 50% batch discount, which each Anthropic model page now shows with the sentence it was read from and the date. The weekly detector that checks the table compares the Standard price of each row against the provider’s page; it does not read the Batch, Flex, Priority or Fast columns, so a change there would go unnoticed until the next reading by hand. Two figures could not be read at all without a browser, xAI’s per-model batch prices and Mistral’s batch column, and this post quotes only what the pages print.

If your workload can wait a day, the Batch prices cut the table’s figures in half at four providers and by a fifth on four xAI models. If it needs an answer in the same session but can live with variable latency and best-effort availability, Flex does the same at OpenAI and Google. If it cannot wait at all, Fast and Priority double them, or nearly, at four: OpenAI, Anthropic on two models, Google at 1.8× and xAI. Either way the number to start from is the Standard one, dated and sourced, in the table.

Frequently asked questions

How much cheaper is the Batch API?

Half price at OpenAI, Anthropic, Google and Mistral, and a fifth off at xAI on Grok 4.3 and the three Grok 4.20 variants, as read on September 16, 2026. The condition is time: OpenAI states "a clear 24-hour turnaround time".

What is the difference between Batch and Flex?

Both cost half of Standard at OpenAI and Google. Batch trades time, with a 24-hour turnaround, while Flex trades capacity: slower responses and occasional resource unavailability at OpenAI, variable latency and best-effort availability at Google.

How much do Priority and Fast mode cost?

Double the Standard price at OpenAI (Fast mode), at Anthropic on Claude Opus 5 and Claude Opus 4.8, and at xAI. Google's Priority tier works out to 1.8× Standard on Gemini 2.5 Pro and Gemini 3.6 Flash, as read on September 16, 2026.

How much cheaper is Muse Spark if Meta can use your data?

Meta's SKU labelled "Used to improve our products" costs $0.10 input, $0.002 cached and $0.20 output, against $1.25, $0.15 and $4.25 for the SKU not used to improve its products: 12.5× less on input. AI Signal publishes the second as the list price.

Written with AI assistance from this site's own verified pricing dataset and reviewed by Álvaro Lucero Rufino, who is responsible for what stays published here. Every figure links to the page where it was verified, with the date it was read; if a number here and a number on that page ever disagree, the page is the one to trust. Automated publishing was discontinued in September 2026; posts that no longer met that standard were removed rather than left up.