Provider pricing
AI provider pricing directory
This directory ranks AI providers by the cheapest chat model each one prices, from frontier labs to inference platforms. Across 20 providers, the lowest blended chat rate is Mistral's Ministral 3B (latest) at $0.040 per 1M tokens. But the cheapest provider for your app depends on which models clear your quality bar and how output-heavy your workload is, not the headline rate. These are list prices; verify the provider source before production billing decisions.
Estimate your cost on any provider - Enter your token mix and volume to get monthly spend →
AI providers ranked by cheapest chat model
Each provider's lowest-cost chat model, blended at a 70% input / 30% output mix from the ByteCosts pricing index, with how many models the provider prices.
| Provider | Cheapest chat model | Input /1M | Output /1M | Blend /1M | Priced models |
|---|---|---|---|---|---|
| Mistral | Ministral 3B (latest) | $0.040 | $0.040 | $0.040 | 28 |
| Cloudflare Workers AI | Granite 4.0 H Micro | $0.017 | $0.112 | $0.045 | 21 |
| Amazon Bedrock | Gemma 3 4B IT | $0.040 | $0.080 | $0.052 | 106 |
| Together AI | LFM2-24B-A2B | $0.030 | $0.120 | $0.057 | 32 |
| Groq | Llama 3.1 8B | $0.050 | $0.080 | $0.059 | 6 |
| Deep Infra | GPT OSS 20B | $0.030 | $0.140 | $0.063 | 39 |
| Cohere | Command R7B | $0.037 | $0.150 | $0.071 | 8 |
| Alibaba | Qwen Turbo | $0.050 | $0.200 | $0.095 | 43 |
| Vertex | GPT OSS 20B | $0.070 | $0.250 | $0.124 | 33 |
| Fireworks AI | GPT OSS 20B | $0.070 | $0.300 | $0.139 | 16 |
| Gemini 2.0 Flash-Lite | $0.075 | $0.300 | $0.142 | 14 | |
| Upstage | solar-mini | $0.150 | $0.150 | $0.150 | 3 |
| OpenAI | GPT-5 Nano | $0.050 | $0.400 | $0.155 | 47 |
| Z.AI | GLM-4.7-FlashX | $0.070 | $0.400 | $0.169 | 12 |
| DeepSeek | DeepSeek Chat | $0.140 | $0.280 | $0.182 | 4 |
| MiniMax (minimax.io) | MiniMax-M2 | $0.300 | $1.20 | $0.570 | 7 |
| Perplexity | Sonar | $1.00 | $1.00 | $1.00 | 4 |
| Moonshot AI | Kimi K2 0711 | $0.600 | $2.50 | $1.17 | 10 |
| xAI | Grok Build 0.1 | $1.00 | $2.00 | $1.30 | 6 |
| Anthropic | Claude Haiku 4.5 | $1.00 | $5.00 | $2.20 | 14 |
Frequently asked questions
Which AI provider is the cheapest?
By cheapest chat model, Mistral leads with Ministral 3B (latest) at $0.040 per 1M tokens blended (70% input / 30% output). Cheapest access rarely means cheapest in production, though: the model still has to clear your quality bar, and output-heavy or retry-heavy workloads can change the ranking, so confirm against your own usage.
How is this different from the provider comparison and pricing pages?
This directory is provider-first: each row is a provider ranked by its cheapest chat model, linking to a full per-provider guide. The compare pages put two frontier flagships head-to-head, and the pricing page ranks the cheapest individual models across every provider.
Where do these provider prices come from?
Each rate comes from the provider's official pricing pages, normalized into the ByteCosts pricing index and dated (updated July 19, 2026). Every row carries a source link and a confidence grade. Prices are list prices and exclude negotiated or volume discounts, so verify the provider source before production billing decisions.
How many models does each provider price?
It varies by provider and is shown next to each one in the directory. A provider needs at least three priced models to get an indexable guide, so long-tail gateways with thin coverage are held back.
AI provider pricing directory. ByteCosts. Updated July 19, 2026. https://bytecosts.com/providers/