Provider pricing
Together AI pricing: API cost per model, user, and month
Together AI lists 32 priced models in the ByteCosts index. This provider page has 11 source-backed model rows, but no source-backed 70% input / 30% output blended chat row to quote, so ByteCosts does not make a lowest-price claim here. For most AI apps the bill is driven by output tokens, retry rate, and prompt-cache hit rate far more than the headline input price, so Together AI is cost-effective when your workload is input-heavy (RAG, classification) or can cache a large shared prefix. Compare Together AI against alternatives on a real workload - seats, requests, and token mix - before committing to a model or a subscription, because a cheaper per-token price can still lose once power users and long outputs are priced in.
Estimate your monthly cost - Turn your usage into Together AI spend →
Together AI pricing at a glance
- 11 source-backed model rows on this provider page; no blended chat minimum is quoted without comparable input/output data
- Last recorded price event: MiniMax M2.7 cache read on Jun 1, 2026
- 32 priced models in the index, updated July 19, 2026
Together AI model prices
Per-million-token list prices for current Together AI models, from the ByteCosts pricing index. Output tokens usually cost several times more than input.
| Model | Input | Output | Context |
|---|---|---|---|
| Juggernaut Pro Flux | $0.0049 | - | - |
| Gemma 4 31B | $0.018 | - | - |
| Multilingual e5 large instruct | $0.020 | - | - |
| Llama 3 8B Instruct Lite | $0.100 | $0.100 | - |
| Qwen3.5 9B | $0.100 | $0.150 | - |
| Qwen3 235B A22B Instruct 2507 FP8 Throughput | $0.200 | $0.600 | - |
| Llama Guard 4 12B | $0.200 | - | - |
| MiniMax M2.7 | $0.300 | $1.20 | - |
| Qwen3.6-Plus | $0.500 | $3.00 | - |
| Kimi K2.6 | $1.20 | $4.50 | - |
| GLM-5.1 | $1.40 | $4.40 | - |
When Together AI is cost-effective
Token price alone does not decide cost. The variables that actually move an AI bill are: how many output tokens each call emits (output is the expensive side), how often calls retry, what fraction of the input prefix can be served from prompt cache, and how heavily your top 1% of users use the product.
Together AI tends to win when your workload is input-heavy with short outputs, when you can cache a large shared system prompt or document context, or when a smaller Together AI model clears your quality bar. It tends to lose when outputs are long and uncached, where a cheaper-per-output model compounds the saving across millions of calls.
Limitations before production billing decisions
Treat ByteCosts calculations as planning estimates, not final billing totals. Real invoices can differ because token mix, retry rate, cache hit rate, rate limits, taxes, gateway fees, regional pricing, and negotiated discounts change the effective cost.
Verify the provider source before production billing decisions, then compare the estimate with your own logs or invoice once production traffic is live.
Frequently asked questions
How much does Together AI cost per million tokens?
Together AI has 32 priced models in the ByteCosts index, and this page has 11 source-backed model rows to inspect. ByteCosts does not have a source-backed 70% input / 30% output low-cost chat blend to quote for this provider page, so it does not make a cheapest-model price claim here. See the price table above for current per-model input/output rates where available.
Is Together AI cheaper than the alternatives?
It depends on the workload. Together AI can be cheaper for input-heavy or cacheable workloads even when its headline price is higher, because output volume and cache hit rate move the bill more than the per-token rate. Use the ByteCosts calculators to compare on your real traffic.
Where does ByteCosts get Together AI prices?
Prices come from Together AI's official pricing/docs pages, normalized into the ByteCosts pricing index and dated. Each record carries a confidence grade and a source link. Prices are list prices and exclude negotiated or volume discounts. Verify the provider source before production billing decisions.
Together AI pricing. ByteCosts. Updated July 19, 2026. https://bytecosts.com/pricing/togetherai/