Provider pricing

Google Vertex AI pricing: API cost per model, user, and month

Google Vertex AI lists 33 priced models in the ByteCosts index, with GPT OSS 20B at about $0.124 per million tokens on ByteCosts' 70% input / 30% output blend. For most AI apps the bill is driven by output tokens, retry rate, and prompt-cache hit rate far more than the headline input price, so Google Vertex AI is cost-effective when your workload is input-heavy (RAG, classification) or can cache a large shared prefix. Compare Google Vertex AI against alternatives on a real workload - seats, requests, and token mix - before committing to a model or a subscription, because a cheaper per-token price can still lose once power users and long outputs are priced in.

Estimate your monthly cost - Turn your usage into Google Vertex AI spend →

Google Vertex AI pricing at a glance

  • Cheapest tracked chat model: GPT OSS 20B at $0.124 per 1M tokens (blended)
  • Flagship: Claude Opus 4.8 at $11.00 per 1M tokens (blended)
  • Last recorded price event: Gemini 3 Pro cached input on Jun 10, 2026
  • 33 priced models in the index, updated July 19, 2026

Google Vertex AI model prices

Per-million-token list prices for current Google Vertex AI models, from the ByteCosts pricing index. Output tokens usually cost several times more than input.

Google Vertex AI model prices per million tokens
ModelInputOutputContextIntelligence
Claude Opus 4.8$5.00$25.001M55.7
Claude Opus 4.7$5.00$25.001M53.5
Claude Sonnet 5$2.00$10.001M53.4
Gemini 3.5 Flash$1.50$9.001M50.2
Claude Sonnet 4.6$3.00$15.001M47.2
Gemini 3.1 Pro Preview$2.00$12.001M46.5
Claude Opus 4.6$5.00$25.001M43.7
Claude Opus 4.5$5.00$25.00200K40.8
GLM-5$1.00$3.20203K39.5
Gemini 3 Flash Preview$0.500$3.001M37.8
GLM-4.7$0.600$2.20200K33.7
Kimi K2 Thinking$0.600$2.50262K32.7

Recent Google Vertex AI price events

Source-backed price events recorded for Google Vertex AI in the ByteCosts ledger, newest first. "New" marks a first-time listing that has no prior value to compare against.

Recent recorded Google Vertex AI price events
ModelRateOldNew%Recorded
Gemini 3 Procached input$0.200new-Jun 10, 2026
Gemini 3 Procache read$0.200new-Jun 10, 2026
Gemini 3 Proimage output-$120.00-Jun 1, 2026
Gemini 3.1 Flashimage output-$60.00-Jun 1, 2026
Gemini 3 Procached input-$0.200-Jun 1, 2026
Gemini 3 Procache read-$0.200-Jun 1, 2026
Gemini 3 Flash Previewcached input-$0.050-Jun 1, 2026
Gemini 3 Flash Previewcache read-$0.050-Jun 1, 2026
Gemini 3.1 Flash-Litecache read-$0.025-Jun 1, 2026
Gemini 3 Flash Previewaudio input-$0.500-Jun 1, 2026
Gemini 3.1 Flash-Liteaudio input-$1.00-Jun 1, 2026
Gemini 2.0 Flash Liteaudio input-$1.00-Jun 1, 2026

When Google Vertex AI is cost-effective

Token price alone does not decide cost. The variables that actually move an AI bill are: how many output tokens each call emits (output is the expensive side), how often calls retry, what fraction of the input prefix can be served from prompt cache, and how heavily your top 1% of users use the product.

Google Vertex AI tends to win when your workload is input-heavy with short outputs, when you can cache a large shared system prompt or document context, or when a smaller Google Vertex AI model clears your quality bar. It tends to lose when outputs are long and uncached, where a cheaper-per-output model compounds the saving across millions of calls.

Hidden costs to watch with Google Vertex AI

List pricing rarely equals your invoice. Budget for these before you commit:

  • Output-token blow-up: reasoning and verbose responses multiply the most expensive token class.
  • Retries and tool loops: a single failed agent step can re-send the whole context.
  • Cache misses: prompt caching only saves money above a break-even reuse count; cold prefixes pay the write premium.
  • Power users: a small share of heavy users can dominate spend and erase a plan's margin.
  • Egress and gateway fees: LLM gateways and observability add a percentage on top of raw model cost.

Limitations before production billing decisions

Treat ByteCosts calculations as planning estimates, not final billing totals. Real invoices can differ because token mix, retry rate, cache hit rate, rate limits, taxes, gateway fees, regional pricing, and negotiated discounts change the effective cost.

Verify the provider source before production billing decisions, then compare the estimate with your own logs or invoice once production traffic is live.

Lower-blend alternatives to Google Vertex AI

On ByteCosts' 70% input / 30% output chat blend, these providers screen lower than Google Vertex AI. Treat this as a shortlist, not a decision - check the selected model against your real token mix:

  • Mistral - $0.040 blended via Ministral 3B (latest), 28 priced models.
  • Cohere - $0.071 blended via Command R7B, 8 priced models.
  • Alibaba Qwen - $0.095 blended via Qwen Turbo, 43 priced models.

Frequently asked questions

How much does Google Vertex AI cost per million tokens?

Google Vertex AI's low-cost tracked chat row is GPT OSS 20B at about $0.124 per million tokens on ByteCosts' 70% input / 30% output blend. That same model's component rates are $0.070 input and $0.250 output; flagship models cost more. See the price table above for current per-model rates.

Is Google Vertex AI cheaper than the alternatives?

It depends on the workload. Google Vertex AI can be cheaper for input-heavy or cacheable workloads even when its headline price is higher, because output volume and cache hit rate move the bill more than the per-token rate. Use the ByteCosts calculators to compare on your real traffic.

Where does ByteCosts get Google Vertex AI prices?

Prices come from Google Vertex AI's official pricing/docs pages, normalized into the ByteCosts pricing index and dated. Each record carries a confidence grade and a source link. Prices are list prices and exclude negotiated or volume discounts. Verify the provider source before production billing decisions.

Google Vertex AI pricing. ByteCosts. Updated July 19, 2026. https://bytecosts.com/pricing/google-vertex/

Sources

Machine-readable