Five Things Decide What You Pay for a Model, and Only One of Them Is the Model
The same weights, or the same first-party API, can cost 20% more, twice as much or ten times as much depending on things that have nothing to do with the model. Two weeks of verifying prices for this site turned up five of them: who serves the model, what time it is, whether a promotion is running, which region you call, and how many tokens the request carries. Each one below is a real case from the pricing table, with the date it was checked and a link to the page where the number can be checked again.
1. The host
Open-weights models are served by several companies, and identical weights do not mean identical prices. On September 6, 2026, Llama 3.3 70B Instruct cost $0.10 per million input tokens on DeepInfra and $1.04 on Together AI: a 10.4× spread for the same model, verified the same day on each host’s own pricing page. DeepSeek V4 Flash 0731 spans 2.75× across its three hosts, from $0.08 on DeepInfra to $0.22 on Fireworks.
That is the exception, not the rule. The homepage on September 15, 2026 put it this way: 6 of the 8 multi-host models in the dataset cost the same on every host, within the 1.15× line the site draws. What separates the two groups is whether the maker sells the model itself with a published list price. GLM-5.3 is $1.40 in and $4.40 out on Together AI and Fireworks AI (verified September 6, 2026) and on Z.ai’s own endpoint on OpenRouter (September 15); at its peak rate, DeepSeek V4 Pro is $1.32 in and $3.96 out on every host that serves it, and the next section is about what happens off-peak. Where nobody sets a reference, every host sets its own. The pattern is laid out in When comparing hosts is worth it, and when it is not.
2. The time of day
DeepSeek V4 Pro lists at $1.32 per million input tokens and $3.96 per million output tokens. That is the rate during DeepSeek’s peak hours, Monday to Friday 01:00 to 04:00 and 06:00 to 10:00 UTC. The rest of the time, every rate is half: $0.66 and $1.98. Peak hours are 35 of the 168 hours in a week, 21%. For 79% of the week the model costs half its list price. The same clock applies to DeepSeek V4.1 Flash: $0.30 and $1.20 at peak, $0.15 and $0.60 off-peak.
Whether a host follows that clock is a separate question. Together AI and Fireworks AI charge $1.32 and $3.96 at all hours. DeepSeek’s own endpoint on OpenRouter follows it: read-only probes on September 11 and 14, 2026, at 02:30, 07:30 and 12:30 UTC, saw $1.32 inside the peak window and $0.66 outside it. This site published the off-peak figure as if it were the list price until September 14, because every earlier check had happened off-peak; the correction is recorded on the model’s page, under its OpenRouter host.
3. The promotion
Gemini 3.6 Flash, 3.7 Flash and 3.8 Flash cost $0.75 per million input tokens and $3.75 per million output tokens, verified September 14, 2026. Those are promotional prices, valid until December 31, 2026. Google prints the list price next to them: $1.50 and $7.50. On January 1, 2027 the bill doubles for anyone who budgeted on the promotional figure. Each of the three pages shows the promotional price, the list price and the date, and records the change as announced, not as a surprise.
The rule this site follows is the same on every page: publish the list price, mark the promotion with its end date, never publish a discount as if it were the price. When a provider publishes only the promotional figure, the page says so and shows no list price rather than a computed one: GPT-5.6 Sol is $4.00 and $20.00, which OpenAI describes as “a 20% reduction in input pricing and a 33% reduction in output pricing”, promotional “at least through November 21, 2026”, with no list figure printed anywhere (checked September 15, 2026). On OpenRouter, where a reseller can undercut the maker at any moment, that rule is the difference between a number you can budget on and a number that changes on its own; the mechanics are in OpenRouter pricing: the number in the catalog isn’t what you pay.
4. The region
Alibaba lists each Qwen model in six regions and two scopes. Beijing and the four regions Alibaba labels Global, Frankfurt, Virginia, Tokyo and Hong Kong, share one price. Singapore, labelled International, is higher. Qwen3.8 Max is $2.00 in and $6.00 out in Singapore against $1.65 and $4.951 elsewhere, 1.21× on input. Qwen3.7 Plus is $0.40 in and $1.60 out in Singapore against $0.276 and $1.101 elsewhere, 1.45×. Both verified September 11, 2026, on the same Alibaba pages.
This site publishes the Singapore figure as the reference, because it is the scope Alibaba calls international and the one every reseller tracked here mirrors, and shows the Beijing and Global figure next to it. Until September 11 the pricing table and the host comparison quoted different regions for the same model; that correction is recorded on each Qwen page too.
5. The size of the request
Several providers charge more once a prompt passes a threshold. GPT-5.6 Sol is $4.00 per million input tokens; above 272,000 tokens of prompt it is $8.00 (verified September 14, 2026), and both figures are the promotional ones from section 3: the long-context tier is defined as 2× input and 1.5× output of the base price, so it moves with it. Gemini 2.5 Pro is $1.25; above 200,000 tokens it is $2.50 (September 14). Qwen3.7 Plus is $0.40 in the Singapore region; above 256,000 tokens it is $1.20 there, three times the headline rate (September 11). Grok 4.6 is $2.00; from 200,000 tokens it is $4.00 (September 14), and xAI’s pricing page states that once a prompt reaches the threshold the long-context rate applies to every token in the request. OpenAI’s model page says the same for Sol, “for the full request”. For Google and Alibaba the pages give the threshold and the rate; how they apply it is not stated there, so it is not stated here.
The counterexample is Anthropic. Claude Opus 4.6 bills its full 1M-token context window at the standard rate, with no long-context surcharge, as its pricing page says and as verified September 14, 2026.
What this means when you compare prices
A headline price is one cell of a table that has five dimensions. Before comparing two models, compare the same cell: same host, same hour, list price rather than promotion, same region, same request size. Every page on this site is built to make that possible: each figure carries the date it was read and a link to the provider’s own page, and where one of the five factors applies, the page says which figure is shown and why. The cost calculator multiplies your token counts by each current model’s headline price, the one in the pricing table: where that price is promotional, as for GPT-5.6 Sol and the three Gemini Flash models today, the estimate is at the promotional rate. It switches to a long-context tier when the average prompt size you enter passes the model’s threshold, and it applies nothing else: no time-of-day tariff and no regional price. If your workload sits on the wrong side of a clock or a region, read the model’s page first.
Frequently asked questions
Why does the same open model cost different amounts on different hosts?
Each host sets its own price when the maker publishes no list price: on September 6, 2026, Llama 3.3 70B Instruct cost $0.10 per million input tokens on DeepInfra and $1.04 on Together AI, a 10.4× spread. Where the maker sells the model with a published list price, hosts tend to match it, and 6 of the 8 multi-host models cost the same on every host on September 15, 2026.
Is DeepSeek cheaper at certain times of day?
Yes. DeepSeek V4 Pro's list price of $1.32 input and $3.96 output per million tokens applies during peak hours, Monday to Friday 01:00 to 04:00 and 06:00 to 10:00 UTC; for the other 79% of the week every rate is half, $0.66 and $1.98.
When does the Gemini Flash promotional price end?
Gemini 3.6 Flash, 3.7 Flash and 3.8 Flash cost $0.75 input and $3.75 output per million tokens until December 31, 2026 (verified September 14, 2026). From January 1, 2027 Google's list price applies: $1.50 and $7.50.
Do long prompts cost more per token?
At several providers, yes: GPT-5.6 Sol's $4.00 input becomes $8.00 above 272,000 tokens, Gemini 2.5 Pro goes from $1.25 to $2.50 above 200,000 and Grok 4.6 from $2.00 to $4.00 from 200,000 (verified September 14, 2026). Anthropic's Claude Opus 4.6 bills its full 1M-token context window at the standard rate.
Written with AI assistance from this site's own verified pricing dataset and reviewed by Álvaro Lucero Rufino, who is responsible for what stays published here. Every figure links to the page where it was verified, with the date it was read; if a number here and a number on that page ever disagree, the page is the one to trust. Automated publishing was discontinued in September 2026; posts that no longer met that standard were removed rather than left up.