Data

When comparing inference hosts is worth it, and when it is not

· Prices verified September 30, 2026 to October 1, 2026

Of the 8 models served by two or more hosts in this dataset, 6 cost practically the same wherever you run them: every host is within 1.15× of the cheapest. Another 2 show a real spread, the widest being Llama 3.3 70B Instruct at 10.4×. For the first group, comparing prices is a waste of your time. For the second, the difference is a multiple, and it is worth ten minutes.

Every price below is a verified host price with a source link; 26 offerings of 13 models across 5 hosts. How this is checked

Most, but not all

Every model here is one that several companies can serve, which is what makes a comparison possible at all; a proprietary model has exactly one seller and belongs in the pricing table, not here. Among these, for some the market has converged and for others it has not. The dataset behind the provider comparison draws the line at 1.15×: if the most expensive host charges no more than 1.15× the cheapest per input token, the price will not decide your choice, and we call that "the same price". Today that line splits the 8 multi-host models into 6 that cost the same and 2 that do not.

Where comparing is not worth it

6 models. Pick on latency, region, rate limits or the tooling you already use; the bill will be the same.

Model Hosts Input Output Spread
DeepSeek V4 Pro 0813 Together AI · OpenRouter $1.32 $3.96 none
GLM-5.3 Together AI · Fireworks AI · OpenRouter $1.40 $4.40 none
gpt-oss-120b Together AI · Fireworks AI · Groq $0.15 $0.60 none
Kimi K3 DeepInfra · Together AI · Fireworks AI · OpenRouter $2.85–$3.00 $14.25–$15.00 1.1×
MiniMax M3 Together AI · Fireworks AI · OpenRouter $0.30 $1.20 none
Qwen3.8 Max Together AI · Fireworks AI $2.00 $6.00 none
DeepSeek V4 Pro 0813: OpenRouter's $1.32 / $3.96 is the maker's list rate, which applies only during the maker's peak hours (Monday to Friday 01:00-04:00 and 06:00-10:00 UTC: 35 of the 168 hours in a week, 21%); the rest of the time that host, like the maker's own API, charges half, $0.66 / $1.98. Together AI charge $1.32 / $3.96 at all hours.

DeepSeek V4 Pro 0813, GLM-5.3, gpt-oss-120b, MiniMax M3 and Qwen3.8 Max cost exactly the same on every host listed. If you are choosing between these hosts on price, you are optimising a rounding error. Spend the time on a latency test from your region instead.

Where it pays to look

2 models, widest spread first. Cheapest and most expensive host for identical weights.

Model Cheapest host Most expensive host Spread
Llama 3.3 70B Instruct DeepInfra $0.10 in · $0.32 out Together AI $1.04 in · $1.04 out 10.4×
DeepSeek V4 Flash 0731 DeepInfra $0.06 in · $0.18 out Together AI $0.14 in · $0.28 out 2.3×

The spread runs from 2.3× on DeepSeek V4 Flash 0731 to 10.4× on Llama 3.3 70B Instruct: at the top end, the most expensive host charges $1.04 per million input tokens for weights another host serves at $0.10. These are the models where a wrong default costs real money every month. Full host-by-host tables, with every price linked to its source, are on the provider comparison.

Three things to keep in mind

A maker's own list price anchors the hosts. Where the maker serves the model itself with a published price, the other hosts sit at or near it (DeepSeek V4 Pro 0813, GLM-5.3, Kimi K3 and MiniMax M3). Both models with a real spread (Llama 3.3 70B Instruct and DeepSeek V4 Flash 0731) lack that anchor: no host is the maker, so each sets its own price. The reverse does not hold: 2 models converge with no maker among the hosts (gpt-oss-120b and Qwen3.8 Max), so the anchor helps, but it is not the only thing that keeps hosts in line.

Spread is on input price. The ratio compares the most and least expensive host per million input tokens, because that is where hosts differ most and what most workloads are dominated by. Output prices usually move with them; the tables show both so you can check for your own mix.

OpenRouter is a router, not a host. Where OpenRouter appears in the tables, the price is the model maker's own list price on OpenRouter's providers page, without promotional discounts. Ask for :floor or sort: "price" and OpenRouter will often find a reseller cheaper than that, with one trade-off: that price changes on its own as promotions end and hosts come and go, so it is a floor for today, not a number to budget on. A default request is spread across several hosts and pays something in between. The mechanics, with evidence, are in OpenRouter pricing: the number in the catalog isn't what you pay. "No list price" on this page means the maker is not among OpenRouter's hosts for that model. It is a different thing from "no published price" on the pricing table (a maker that sells the model but publishes no per-token price for it: today Meta for 1 current model and Cohere for 4 current models) and from "no data" (a price we could not verify).

A further 7 models in the dataset have a single host; for 2 of them there is no list price to show. There is nothing to compare there; the only question is whether you accept a single vendor for that model.

Verification

All figures on this page are computed at build time from the same dataset and the same rules as the provider comparison, so the two pages cannot disagree. An offering is included only if it carries a verification date and a link to the host's official pricing page; verification dates range from September 30, 2026 to October 1, 2026. The 1.15× threshold is a fixed editorial cut, not a measurement; move it and the split moves with it. Prices are USD per million tokens. Nothing is estimated.