Data study
The same AI model can cost 10x more depending on the provider
ByteCosts tracks 57 AI models offered by two or more providers with a comparable per-million-token list price. Most cross-listed models are priced identically across hosts, but 10 carry a 50%-or-more spread and 8 cost at least double. The widest spread is Kimi K2.6: 956% (10x) between the cheapest listing (novita-ai, $0.540 combined) and the priciest (togetherai, $5.70 combined). These are list prices normalized to USD per 1M input plus 1M output tokens; the same model name can differ across hosts by quantization, region, and hosting overhead, so verify the provider source before switching or budgeting.
Open the pricing index - Inspect every provider’s source and confidence grade →
What the data shows
Computed from the committed ByteCosts pricing index at build time:
- 57 models are listed by at least two providers with both an input and an output price.
- The median cross-provider spread is 0%most resellers price a given model at the same list rate.
- 10 models have a spread of 50% or more, and 8 cost at least double on the priciest provider.
- The widest gap is Kimi K2.6 at 956% (10x): novita-ai $0.540 vs togetherai $5.70 for 1M input plus 1M output tokens.
Biggest provider price spreads
Models with the largest gap between the cheapest and priciest provider listing, ranked by premium. Combined price is 1M input plus 1M output tokens.
| Model | Providers | Cheapest listing | Priciest listing | Premium |
|---|---|---|---|---|
| Kimi K2.6 | 5 | novita-ai · $0.540 | togetherai · $5.70 | 956% |
| DeepSeek-V4-Flash | 4 | deepinfra · $0.300 | deepseek · $1.76 | 487% |
| Qwen3-Max | 2 | alibaba · $1.47 | deepinfra · $7.20 | 390% |
| DeepSeek-V3.2 | 3 | deepinfra · $0.640 | friendli · $2.00 | 212% |
| openai/gpt-oss-120b | 2 | novita-ai · $0.300 | groq · $0.750 | 150% |
| DeepSeek V3.1 | 2 | deepinfra · $1.00 | baseten · $2.00 | 100% |
| Gemini 3 Flash Preview | 3 | google-vertex · $1.75 | vivgrid · $3.50 | 100% |
| MiniMax-M3 | 4 | minimax · $1.50 | opencode-go · $3.00 | 100% |
| openai/gpt-oss-20b | 2 | novita-ai · $0.190 | groq · $0.375 | 97% |
| deepseek-ai/DeepSeek-V4-Pro | 3 | siliconflow · $1.29 | vultr · $2.20 | 71% |
| GLM 5.1 | 7 | lilac · $3.90 | zhipuai · $5.80 | 49% |
| DeepSeek-V4-Pro | 4 | deepinfra · $3.90 | deepseek · $5.28 | 35% |
| Minimax M2 | 2 | novita-ai · $1.50 | amazon-bedrock · $1.81 | 21% |
| Minimax M2.1 | 2 | novita-ai · $1.50 | amazon-bedrock · $1.80 | 20% |
| Minimax M2.5 | 4 | baseten · $1.50 | amazon-bedrock · $1.80 | 20% |
| Claude Sonnet 4.6 | 3 | anthropic · $18.00 | google-vertex-anthropic · $19.80 | 10% |
| Claude Sonnet 4.5 | 2 | anthropic · $18.00 | google-vertex-anthropic · $19.80 | 10% |
| Claude Opus 4.6 | 2 | anthropic · $30.00 | google-vertex-anthropic · $33.00 | 10% |
| Claude Opus 4.5 | 2 | anthropic · $30.00 | google-vertex-anthropic · $33.00 | 10% |
| Gemini 3.1 Pro Preview | 2 | google · $14.00 | vivgrid · $15.00 | 7% |
How the spread is measured
Models are matched across providers by a normalized name key (lowercase, punctuation-stripped), so "GPT-OSS-120B", "gpt-oss-120b" and "openai/gpt-oss-120b" resolve to the same model.
Each provider contributes its cheapest committed listing for that model, so region and tier variants never inflate the premium. The combined price is the sum of the per-1M-token input and output list prices.
The premium is the priciest listing divided by the cheapest, minus one. A 0% spread means every provider lists the same combined rate.
Only records with both an input and an output price and a non-zero combined total are included, so free tiers and embedding/audio-only rows do not distort the ranking.
What the spread does not mean
- A wider gap is not necessarily gouging: a “same-named” model can differ by quantization, context window, region, throughput, or a bundled gateway/SLA.
- List prices exclude negotiated and committed-use discounts, taxes, and gateway fees, the price you actually pay can differ from the listed rate.
- The data is a committed snapshot and refreshes only after a manual refresh; verify the provider source (linked on each pricing page) before acting.
- A cheaper listing is not automatically the better deal: latency, uptime, and rate limits are not captured in token list prices.
Frequently asked questions
Why does the same AI model cost different prices on different providers?
The same model name can be hosted with different quantization, context windows, regions, or service tiers, and resellers add their own margin and infrastructure costs. A small number also list genuinely different rates for the same artifact, which is what this page surfaces.
Is the cheapest listing always the best choice?
No. Token list price is one signal among several. Latency, uptime, rate limits, data-residency, and the provider’s terms all count, and those are not captured in a per-token list price. Use the spread as a starting point, then verify the provider source.
Are these live prices?
No. This page is computed from the committed ByteCosts pricing index generated August 22, 2026, not a live provider API. Prices change, so re-check the last-checked date on each provider page before budgeting.
AI model price spread across providers. ByteCosts. Updated August 22, 2026. https://bytecosts.com/model-price-spread/