Self-host an open model
Self-host MiniMax M2.5: GPU, VRAM, and rental cost
Self-hosting MiniMax M2.5 (228.7B) in FP8 needs about 321 GB of VRAM per GPU at 8K context and 8 concurrent requests, which exceeds every single tracked GPU, so it requires tensor parallelism across multiple GPUs. Lower the context, concurrency, or precision to fit fewer cards.
Estimate cost per 1M tokens - Self-host MiniMax M2.5 serving cost →
Reference workload for self-hosting MiniMax M2.5
No tracked single GPU holds MiniMax M2.5 at this reference shape, so the page shows the VRAM inputs instead of a rental table.
| Input | Value |
|---|---|
| Precision | FP8 |
| Context | 8,192 tokens |
| Concurrency | 8 sequences |
| Required VRAM per GPU | about 321 GB |
| Largest tracked GPU | GB300 (Grace Blackwell Ultra, per GPU) (288 GB) |
Frequently asked questions
What GPU do I need to run MiniMax M2.5?
MiniMax M2.5 needs about 321 GB of VRAM per GPU in FP8 at 8K context, which is more than any single tracked GPU, so it requires splitting across multiple GPUs (tensor parallelism).
How is the VRAM figure calculated?
Model weights (parameters times bytes per weight for the precision) plus the KV cache (from the model’s real layers, KV heads, head dimension, and attention pattern) plus activation and a safety margin. Architecture comes from the model’s Hugging Face config; GPU VRAM from the NVIDIA datasheet.
Self-host MiniMax M2.5: GPU and VRAM. ByteCosts. Updated October 10, 2026. https://bytecosts.com/gpu/self-host/minimax-m2-5/