ByteCosts
serverless-gpu-vs-always-on.calc
Estimate

GPU economics

Serverless GPU vs Always-On Calculator

Compare allocation-session billing, cold starts, warm idle and minimum charges with an equivalent always-on GPU fleet. Low traffic can favor an allocation-based service because an always-on GPU incurs charges while idle. Cold starts, minimum allocation periods and warm-idle policies can reverse the apparent saving. The calculator bills allocation sessions rather than pretending every API request creates a new container. An aggregate worker-hour check catches impossible always-on comparisons, but it does not prove that a smaller fleet can satisfy bursts. Compare the same hardware, workload and latency target before treating the cost difference as actionable.

Open the live Serverless vs Always-On calculator - Serverless vs Always-On →

Example scenario

Illustrative example: 1,000 allocation sessions with 60 seconds of work each include 100 cold starts adding 10 seconds. At a one-second increment and 60-second minimum, billable time is 61,000 seconds. A hypothetical rate of 0.001 dollars per second costs 61 dollars before other charges. One always-on replica at 1 dollar per hour costs 730 dollars over a 730-hour planning period, but it must still meet the same arrival pattern and latency target.

What the inputs mean

  • Allocation sessions: Enter sessions for the same workload and billing period. The initial value is illustrative, not an observed provider rate.
  • Sessions billed a cold start: Enter sessions for the same workload and billing period. The initial value is illustrative, not an observed provider rate.
  • Active work in each session: Enter seconds for the same workload and billing period. The initial value is illustrative, not an observed provider rate.
  • Billed overhead per cold session: Enter seconds for the same workload and billing period. The initial value is illustrative, not an observed provider rate.
  • Minimum allocation bill: Enter seconds for the same workload and billing period. The initial value is illustrative, not an observed provider rate.
  • Billing increment: Enter seconds for the same workload and billing period. The initial value is illustrative, not an observed provider rate.
  • Additional billed warm idle: Enter seconds for the same workload and billing period. The initial value is illustrative, not an observed provider rate.
  • Serverless allocation rate: Enter USD / second for the same workload and billing period. The initial value is illustrative, not an observed provider rate.
  • Other serverless costs: Enter USD for the same workload and billing period. The initial value is illustrative, not an observed provider rate.
  • Comparison period: Enter hours for the same workload and billing period. The initial value is illustrative, not an observed provider rate.
  • Always-on replicas: Enter replicas for the same workload and billing period. The initial value is illustrative, not an observed provider rate.
  • Always-on replica rate: Enter USD / hour for the same workload and billing period. The initial value is illustrative, not an observed provider rate.
  • Other always-on costs: Enter USD for the same workload and billing period. The initial value is illustrative, not an observed provider rate.

What the result means

The result exposes the cost components and capacity constraint that determine this scenario. A missing feasible crossover is shown explicitly rather than converted into a zero-cost result.

Assumptions

  • Low traffic can favor an allocation-based service because an always-on GPU incurs charges while idle. Cold starts, minimum allocation periods and warm-idle policies can reverse the apparent saving. The calculator bills allocation sessions rather than pretending every API request creates a new container. An aggregate worker-hour check catches impossible always-on comparisons, but it does not prove that a smaller fleet can satisfy bursts. Compare the same hardware, workload and latency target before treating the cost difference as actionable.
  • Default rates and workloads are original hypothetical examples. No measured provider performance or negotiated price is implied.
  • Monthly conversion hours are a planning input, not a universal billing rule. GB and GiB must be normalized before entering storage or traffic quantities.
  • Copied results contain numeric assumptions and scenario explanations only. Model weights, prompts, API keys and private invoices are never requested.

Where the prices come from

This calculator does not automatically fill prices from a provider catalog. Verify the exact SKU, billing unit, region, tax treatment and effective date, then enter the applicable rate. Unknown charges remain unresolved until supplied.

Formula and methodology

Billed seconds = warm sessions x rounded(max(active seconds, minimum)) + cold sessions x rounded(max(active + cold-start seconds, minimum)) + billed warm idle. Multiply by the per-second rate and add other costs. Always-on cost = replicas x period hours x hourly rate + other costs.

Interpretation guide

  • Compare alternatives with the same workload assumptions.
  • Stress-test output-heavy, retry-heavy, cache-miss, and power-user cases before committing budget.
  • Verify source links and production logs before using the estimate for billing decisions.

Limitations before production billing decisions

Treat ByteCosts calculations as planning estimates, not final billing totals. Real invoices can differ because token mix, retry rate, cache hit rate, rate limits, taxes, gateway fees, regional pricing, and negotiated discounts change the effective cost.

Verify the provider source before production billing decisions, then compare the estimate with your own logs or invoice once production traffic is live.

Serverless GPU vs Always-On Calculator. ByteCosts. https://bytecosts.com/tools/serverless-gpu-vs-always-on/

Sources

USD per 1M tokenslist prices, excl. discounts & taxupdated 2026-10-09Cite this data