LLM API pricing / Alibaba
What the Qwen API costs, model by model
Alibaba sells 6 current models with a verified price on its API, from $0.03 to $2.50 per 1M input tokens and from $0.13 to $7.50 per 1M output tokens. Every price below was read on a page Alibaba publishes, verified between September 11, 2026 and September 28, 2026, and links to it.
Current models
| Model | Input $/1M | Output $/1M | Cached input | Context | Verified |
|---|---|---|---|---|---|
| Qwen3.7-Flash | $0.03 | $0.13 | $0.006 | 1M | 2026-09-17 |
| Qwen3.8-Flash | $0.15 | $0.47 | $0.016 | 1M | 2026-09-11 |
| Qwen3.8-Omni-Flash | $0.15 | $0.47 | $0.016 | 1M | 2026-09-28 |
| Qwen3.7-Plus | $0.40 | $1.60 | $0.08 | 1M | 2026-09-11 |
| Qwen3.8-Max | $2.00 | $6.00 | $0.25 | 1M | 2026-09-11 |
| Qwen3.7-Max | $2.50 | $7.50 | $0.50 | 1M | 2026-09-11 |
How Alibaba prices its models
- 6 of the 6 models publish a cached-input price, at 10.7%, 12.5% or 20% of the input price.
- 2 models charge more for long prompts : Qwen3.7-Flash $0.10 / $0.40 for prompts over 32,000 tokens, Qwen3.7-Plus $1.20 / $4.80 for prompts over 256,000 tokens.
- Prices on this site are Alibaba's Singapore (International) prices; in Beijing and Alibaba's Global regions (Frankfurt, Virginia, Tokyo, Hong Kong) the same models cost: Qwen3.7-Flash $0.028 / $0.11, Qwen3.8-Flash $0.113 / $0.382, Qwen3.8-Omni-Flash $0.113 / $0.382, Qwen3.7-Plus $0.276 / $1.101, Qwen3.8-Max $1.65 / $4.951, Qwen3.7-Max $1.65 / $4.951.
Recorded price changes
- September 17, 2026 · added to this dataset, our omission. Omission on our side, not a price change by Alibaba. Qwen3.7 Flash belongs to the same generation as the Qwen3.7 Plus and Max rows and had no row in this dataset. It turned up on September 17, 2026 through the OpenRouter signal of this site's read-only scan for missing models, and its prices were read that day on Alibaba's own Model Studio page: $0.03 input, $0.13 output and $0.006 cached per million tokens in the Singapore (International) region for prompts up to 32K tokens, $0.10 / $0.40 above 32K and $0.20 / $0.80 above 256K. For Alibaba the table covers the current Qwen generations, 3.7 and 3.8; the older generations Alibaba still sells, found the same day, were left out on purpose. source
- September 11, 2026 · change of this site's criterion. Change of criterion on our side, not a price change by Alibaba: no Alibaba price rose. Alibaba Cloud's Model Studio pages list the same model at two price levels: the China (Beijing) region and the Global regions (Frankfurt, Virginia, Tokyo, Hong Kong) share one level, and the Singapore region, labelled by Alibaba "Scope: International", is higher. Until September 11, 2026 the first-party table quoted the Beijing/Global level while the provider comparison quoted Alibaba's OpenRouter endpoint at the Singapore level, so the site showed two different "Alibaba" prices for the same model. From September 11 the reference region is Singapore (International), the one Alibaba labels international, the one every reseller we track (OpenRouter, Together, Fireworks) mirrors, and the higher of the two; the Beijing/Global figures stay published next to it. Published figures move from 1.65 / 4.951 to 2.00 / 6.00 (Qwen3.8 Max), 0.113 / 0.382 to 0.15 / 0.47 (Qwen3.8 Flash), 0.276 / 1.101 to 0.40 / 1.60 (Qwen3.7 Plus, prompts up to 256K; the tier above 256K, 1.20 / 4.80, is now recorded too) and 1.65 / 4.951 to 2.50 / 7.50 (Qwen3.7 Max). Alibaba's pages state that they show the original pricing excluding any limited-time promotions. source
Analysis that cites Alibaba models
- Prompt Caching Is Not 10% Everywhere: What a Cache Hit Costs Across Eight Providers · names 5 Alibaba models
- Five Things Decide What You Pay for a Model, and Only One of Them Is the Model · names 2 Alibaba models
- Same List Price, Different Bill: Five Pairs of LLM APIs Where the Workload Decides What You Pay · names 1 Alibaba model
How every figure is checked: methodology. All providers: the full pricing table.