ByteCosts
qwen3-5-9b.board
Snapshot

Self-host an open model

Self-host Qwen3.5 9B: GPU, VRAM, and rental cost

Self-hosting Qwen3.5 9B (9.7B) in FP8 needs about 16 GB of VRAM per GPU at 8K context and 8 concurrent requests. It fits on 21 tracked GPUs on a single card, the cheapest being the L40 48GB from about $0.20 per GPU-hour. Larger context or higher concurrency grows the KV cache and can push it past one card into tensor parallelism. This is a fit and rental-rate estimate, not a throughput quote; use the calculators below for cost per token.

Estimate cost per 1M tokens - Self-host Qwen3.5 9B serving cost →

GPUs that hold Qwen3.5 9B on one card

Single-GPU fit in FP8 at 8K context, 8 concurrent requests, with the cheapest tracked rental rate.

Qwen3.5 9B single-GPU fit and cheapest rental
GPUVRAMCheapest /GPU-hrProvider
L40 48GB48 GB$0.20Massed Compute
H100 NVL 94GB94 GB$0.36Massed Compute
RTX PRO 6000 96GB96 GB$0.55Massed Compute
RTX 6000 Ada 48GB48 GB$0.57Massed Compute
L4 24GB24 GB$0.80Modal
L40S 48GB48 GB$0.97Massed Compute
A10 24GB24 GB$1.10Modal
A100 SXM 40GB40 GB$1.29Verda
A100 PCIe 80GB80 GB$1.35Massed Compute
A100 SXM 80GB80 GB$1.38Massed Compute
H100 SXM 80GB80 GB$1.99Voltage Park
AMD Instinct MI300X 192GB192 GB$3.45Crusoe
H200 SXM 141GB141 GB$3.62Massed Compute
B200 (HGX, per GPU)180 GB$5.43Massed Compute
GH200 Grace Hopper 96GB HBM396 GB$6.50CoreWeave
B300 (Blackwell Ultra, per GPU)288 GB$6.60Massed Compute
GB300 (Grace Blackwell Ultra, per GPU)288 GB$8.62Verda
GB200 (Grace Blackwell, per GPU)186 GB$10.50CoreWeave
A10G 24GB (AWS)24 GBn/auntracked
V100 PCIe 32GB32 GBn/auntracked
Quadro RTX 6000 24GB24 GBn/auntracked

Frequently asked questions

What GPU do I need to run Qwen3.5 9B?

In FP8 at 8K context, Qwen3.5 9B needs about 16 GB of VRAM per GPU. The cheapest single GPU that holds it is the L40 48GB (48 GB) from around $0.20 per GPU-hour. Higher context or concurrency needs more VRAM or tensor parallelism.

How is the VRAM figure calculated?

Model weights (parameters times bytes per weight for the precision) plus the KV cache (from the model’s real layers, KV heads, head dimension, and attention pattern) plus activation and a safety margin. Architecture comes from the model’s Hugging Face config; GPU VRAM from the NVIDIA datasheet.

Self-host Qwen3.5 9B: GPU and VRAM. ByteCosts. Updated September 6, 2026. https://bytecosts.com/gpu/self-host/qwen3-5-9b/

Sources

USD / GPU-hourlist prices, excl. discounts & taxupdated 2026-09-04Cite this data