Self-host an open model
Self-host Gemma 4 31B: GPU, VRAM, and rental cost
Self-hosting Gemma 4 31B (31.3B) in FP8 needs about 64 GB of VRAM per GPU at 8K context and 8 concurrent requests. It fits on 12 tracked GPUs on a single card, the cheapest being the H100 NVL 94GB from about $0.36 per GPU-hour. Larger context or higher concurrency grows the KV cache and can push it past one card into tensor parallelism. This is a fit and rental-rate estimate, not a throughput quote; use the calculators below for cost per token.
Estimate cost per 1M tokens - Self-host Gemma 4 31B serving cost →
GPUs that hold Gemma 4 31B on one card
Single-GPU fit in FP8 at 8K context, 8 concurrent requests, with the cheapest tracked rental rate.
| GPU | VRAM | Cheapest /GPU-hr | Provider |
|---|---|---|---|
| H100 NVL 94GB | 94 GB | $0.36 | Massed Compute |
| RTX PRO 6000 96GB | 96 GB | $0.55 | Massed Compute |
| A100 PCIe 80GB | 80 GB | $1.35 | Massed Compute |
| A100 SXM 80GB | 80 GB | $1.38 | Massed Compute |
| H100 SXM 80GB | 80 GB | $1.99 | Voltage Park |
| AMD Instinct MI300X 192GB | 192 GB | $3.45 | Crusoe |
| H200 SXM 141GB | 141 GB | $3.62 | Massed Compute |
| B200 (HGX, per GPU) | 180 GB | $5.43 | Massed Compute |
| GH200 Grace Hopper 96GB HBM3 | 96 GB | $6.50 | CoreWeave |
| B300 (Blackwell Ultra, per GPU) | 288 GB | $6.60 | Massed Compute |
| GB300 (Grace Blackwell Ultra, per GPU) | 288 GB | $8.62 | Verda |
| GB200 (Grace Blackwell, per GPU) | 186 GB | $10.50 | CoreWeave |
Frequently asked questions
What GPU do I need to run Gemma 4 31B?
In FP8 at 8K context, Gemma 4 31B needs about 64 GB of VRAM per GPU. The cheapest single GPU that holds it is the H100 NVL 94GB (94 GB) from around $0.36 per GPU-hour. Higher context or concurrency needs more VRAM or tensor parallelism.
How is the VRAM figure calculated?
Model weights (parameters times bytes per weight for the precision) plus the KV cache (from the model’s real layers, KV heads, head dimension, and attention pattern) plus activation and a safety margin. Architecture comes from the model’s Hugging Face config; GPU VRAM from the NVIDIA datasheet.
Self-host Gemma 4 31B: GPU and VRAM. ByteCosts. Updated September 6, 2026. https://bytecosts.com/gpu/self-host/gemma-4-31b/