GPU calculator

Self-host LLM cost per 1M tokens calculator

Self-host LLM cost per 1M tokens calculator is built for teams weighing self-hosting an open model against a hosted API. Use it to decide what a self-hosted open model costs per 1M tokens on a rented GPU at real utilization. Keep the workload assumptions consistent across options, then inspect the cited prices and last-checked dates before committing budget.

Open the cost calculator - Self-host cost per 1M tokens →

The decision this page helps you make

Estimate the cost per 1M tokens to self-host an open model on a rented GPU, from the GPU rental rate, utilization, and serving throughput, with a labeled band when throughput is not measured.

The practical question is what a self-hosted open model costs per 1M tokens on a rented GPU at real utilization. Use the same workload assumptions for every option so the comparison reflects billing differences instead of different inputs.

Start with these inputs

  • GPU rate: Cheapest per-GPU-hour from the pricing index.
  • Throughput: Measured benchmark, or a labeled roofline band.
  • Output: Cost per 1M tokens, range, and monthly run-rate.

What the result includes

AreaWhat ByteCosts shows
GPU rateCheapest per-GPU-hour from the pricing index
ThroughputMeasured benchmark, or a labeled roofline band
OutputCost per 1M tokens, range, and monthly run-rate

How to use the result

  • Run a realistic base case and a heavier-usage case before choosing a provider or plan.
  • Compare alternatives with identical traffic, token, seat, runtime, and retry assumptions.
  • Open the cited provider source before a purchase or production billing decision.

Formula

monthlyCost = usageVolume * unitCost, adjusted only for the billing units and optional inputs that this calculator exposes.

Assumptions

  • Published rates come from committed ByteCosts datasets or visible source-backed rows.
  • Calculator outputs are planning estimates, not final invoices.
  • Taxes, negotiated discounts, billing minimums, and undocumented limits are excluded unless the page states otherwise.
  • Unknown inputs stay unknown until the user supplies them or a source-backed value is available.

Example scenario

Enter a conservative base case, then duplicate it and change one important driver such as usage, retries, utilization, or output volume. Comparing controlled scenarios makes the result easier to explain and audit.

Interpretation guide

  • Compare alternatives with identical workload assumptions.
  • Stress-test the input that is most likely to grow in production.
  • Verify source links and last-checked dates before making a purchase decision.

Limitations

Self-host LLM cost per 1M tokens calculator is a planning tool, not a billing guarantee. It uses the visible assumptions and committed source-backed data available at the page's last update.

Check the cited provider page and your own production logs before signing a contract, changing price, or committing infrastructure spend.

Frequently asked questions

What should I enter first in Self-host LLM cost per 1M tokens calculator?

Start with gpu rate: cheapest per-gpu-hour from the pricing index. Add optional adjustments only after the base case is understandable.

Is the result a guaranteed invoice forecast?

No. It is a planning estimate based on the visible workload assumptions and source-backed public prices. Taxes, negotiated discounts, undocumented limits, and production behavior can change the final invoice.

Where do the prices and assumptions come from?

ByteCosts keeps provider source links, confidence information, and last-checked dates attached to pricing records. User-entered workload assumptions remain separate from published vendor facts.

Self-host LLM cost per 1M tokens calculator. ByteCosts. https://bytecosts.com/tools/open-model-token-cost/

Sources

Machine-readable