ByteCosts
groq.table
Live

Provider pricing

Groq pricing: API cost per model, user, and month

Groq lists 5 priced models in the ByteCosts index. This provider page has 10 source-backed model rows, but no source-backed 70% input / 30% output blended chat row to quote, so ByteCosts does not make a lowest-price claim here. For most AI apps the bill is driven by output tokens, retry rate, and prompt-cache hit rate far more than the headline input price, so Groq is cost-effective when your workload is input-heavy (RAG, classification) or can cache a large shared prefix. Compare Groq against alternatives on a real workload, seats, requests, and token mix, before committing to a model or a subscription, because a cheaper per-token price can still lose once power users and long outputs are priced in.

Estimate your monthly cost - Turn your usage into Groq spend →

Groq pricing at a glance

  • 10 source-backed model rows on this provider page; no blended chat minimum is quoted without comparable input/output data
  • Last recorded price event: GPT OSS 20B 128k batch input on Aug 22, 2026
  • 5 priced models in the index, updated August 22, 2026

Groq model prices

Per-million-token list prices for current Groq models, from the ByteCosts pricing index. Output tokens usually cost several times more than input.

Groq model prices per million tokens
ModelInputOutputContext
Llama 3.1 8B Instant 128k$0.050$0.080128K
GPT OSS 20B 128k$0.075$0.300128K
GPT OSS Safeguard 20B$0.075$0.300128K
openai/gpt-oss-20b$0.075$0.300-
Llama 4 Scout (17Bx16E) 128k$0.110$0.340128K
GPT OSS 120B 128k$0.150$0.600128K
openai/gpt-oss-120b$0.150$0.600-
Qwen3 32B 131k$0.290$0.590131K
Llama 3.3 70B Versatile 128k$0.590$0.790128K
moonshotai/kimi-k2-instruct-0905$1.00$3.00-

Recent Groq price events

Source-backed price events recorded for Groq in the ByteCosts ledger, newest first. Every event below is a first-time listing, the ledger started tracking that rate; the model itself may be older.

Recent recorded Groq price events
ModelRatePriceRecorded
GPT OSS 20B 128kbatch input$0.037Aug 22, 2026
GPT OSS 20B 128kbatch output$0.150Aug 22, 2026
GPT OSS Safeguard 20Bbatch input$0.037Aug 22, 2026
GPT OSS Safeguard 20Bbatch output$0.150Aug 22, 2026
GPT OSS 120B 128kbatch input$0.075Aug 22, 2026
GPT OSS 120B 128kbatch output$0.300Aug 22, 2026
Llama 4 Scout (17Bx16E) 128kbatch input$0.055Aug 22, 2026
Llama 4 Scout (17Bx16E) 128kbatch output$0.170Aug 22, 2026
Qwen3 32B 131kbatch input$0.145Aug 22, 2026
Qwen3 32B 131kbatch output$0.295Aug 22, 2026
Llama 3.3 70B Versatile 128kbatch input$0.295Aug 22, 2026
Llama 3.3 70B Versatile 128kbatch output$0.395Aug 22, 2026

When Groq is cost-effective

Token price alone does not decide cost. The variables that move an AI bill are: how many output tokens each call emits (output is the expensive side), how often calls retry, what fraction of the input prefix can be served from prompt cache, and how heavily your top 1% of users use the product.

Groq tends to win when your workload is input-heavy with short outputs, when you can cache a large shared system prompt or document context, or when a smaller Groq model clears your quality bar. It tends to lose when outputs are long and uncached, where a cheaper-per-output model compounds the saving across millions of calls.

Hidden costs to watch with Groq

List pricing rarely equals your invoice. Budget for these before you commit:

  • Output-token blow-up: reasoning and verbose responses multiply the most expensive token class.
  • Retries and tool loops: a single failed agent step can re-send the whole context.
  • Cache misses: prompt caching only saves money above a break-even reuse count; cold prefixes pay the write premium.
  • Power users: a small share of heavy users can dominate spend and erase a plan’s margin.
  • Egress and gateway fees: LLM gateways and observability add a percentage on top of raw model cost.

Limitations before production billing decisions

Treat ByteCosts calculations as planning estimates, not final billing totals. Real invoices can differ because token mix, retry rate, cache hit rate, rate limits, taxes, gateway fees, regional pricing, and negotiated discounts change the effective cost.

Verify the provider source before production billing decisions, then compare the estimate with your own logs or invoice once production traffic is live.

Frequently asked questions

How much does Groq cost per million tokens?

Groq has 5 priced models in the ByteCosts index, and this page has 10 source-backed model rows to inspect. ByteCosts does not have a source-backed 70% input / 30% output low-cost chat blend to quote for this provider page, so it does not make a cheapest-model price claim here. See the price table above for current per-model input/output rates where available.

Is Groq cheaper than the alternatives?

It depends on the workload. Groq can be cheaper for input-heavy or cacheable workloads even when its headline price is higher, because output volume and cache hit rate move the bill more than the per-token rate. Use the ByteCosts calculators to compare on your real traffic.

Where does ByteCosts get Groq prices?

Prices come from Groq‘s official pricing/docs pages, normalized into the ByteCosts pricing index and dated. Each record carries a confidence grade and a source link. Prices are list prices and exclude negotiated or volume discounts. Verify the provider source before production billing decisions.

Groq pricing. ByteCosts. Updated August 22, 2026. https://bytecosts.com/pricing/groq/

Sources

Machine-readable

USD per 1M tokenslist prices, excl. discounts & taxupdated 2026-08-22Cite this data