viabandwidthGPU Compute
viabandwidthBuyer's Guide

On-Demand GPU Pricing

On-demand, spot and committed are a gradient of price against certainty. This is what on-demand really means, how the three modes differ, and when paying on-demand is the right call.

Sponsor this paper →
1,009
GPU providers indexed
19
Accelerator models tracked
23
Countries with indexed capacity
144
Network-verified direct operators

The short answer

On-demand means you pay per unit of time for a GPU instance with no commitment, spin it up when you need it and stop paying when you stop it. That flexibility is the most expensive way to buy a GPU hour and it is the right way for work that is variable or not yet understood. Spot capacity sits below it in price and above it in risk, and committed terms sit below it in price and above it in obligation. Reading a rate correctly means knowing which of the three you are looking at.

01 · The modeWhat on-demand actually means

On-demand pricing is the pay-as-you-go rate for a GPU instance. You start it when you want it, you are billed by the hour or the second while it runs and the meter stops when you stop the machine, with no term to sign and no capacity reserved in advance. That is the model that lets you scale with demand and walk away cleanly.

The convenience carries a premium. Because the provider is holding capacity ready for you to claim at any moment and asking nothing in return, the on-demand rate is the highest of the standard pricing modes. A workload that runs almost all the time on on-demand is quietly paying for a flexibility it never uses, which is the single most common way GPU bills end up larger than they needed to be.

02 · The gradientOn-demand, spot and committed

Think of the three modes as a gradient of price against certainty. Spot is the cheapest, because you are renting hardware the operator would otherwise leave idle and accepting that it can be reclaimed with little notice, which suits interruptible batch work that checkpoints and resumes. Committed terms are also cheaper than on-demand, because you are promising to pay for a fixed slice of capacity over a year or more. On-demand sits in the middle on price and at the top on flexibility, so almost no efficient setup uses one mode for everything.

03 · RegionWhere on-demand capacity sits

The accelerator model is the largest lever on an on-demand rate and region is the next, because capacity is priced differently around the world and a cheaper region far from your users can cost you in latency what it saves in rate. The counts below are the number of tracked providers with indexed capacity in each country.

United States62Netherlands48Germany43India39United Kingdom37Singapore31Canada26Japan23TRACKED GPU PROVIDERS, BY COUNTRY
Tracked GPU providers by country. Breadth outside the United States gives on-demand buyers real choice of region for latency and residency.
CountryProviders indexed
United States62
Netherlands48
Germany43
India39
United Kingdom37
Singapore31
Canada26
Japan23

A cheaper on-demand rate in a distant region is only cheaper until latency or egress is counted, so treat geography as a filter on the shortlist rather than a way to chase the lowest number.

What the on-demand catalog actually offers

The accelerator model is the largest single lever on an on-demand rate, and the depth of supply behind each card decides how much that rate can be shopped. The chart below is the number of tracked providers listing each accelerator.

H100224A100181H200147B200102L40S101RTX 409076V10064A4049PROVIDERS LISTING EACH ACCELERATOR
Providers listing each accelerator. The deep prior-generation supply is where on-demand competition is fiercest and the premium over spot is smallest.
AcceleratorProviders listing it
H100224
A100181
H200147
B200102
L40S101
RTX 409076
V10064
A4049

On-demand rates move most on the cards with the deepest supply, so a workload that can run on an A100 or an L40S will find the on-demand premium over spot far easier to swallow than one that needs the scarce frontier silicon.

04 · The callWhen on-demand is the right choice

  • 01Use on-demand while the workload is variable or unknown. Prototyping, model evaluation and genuinely spiky traffic are what the premium buys.
  • 02Move the steady part off on-demand. Once the shape settles, a committed floor or a dedicated server is cheaper for the predictable work.
  • 03Read what the rate includes. Whole-machine flat pricing and accelerator-only pricing with separate storage and egress can diverge sharply on a real month.
  • 04Price a month, not an hour. Turn the per-hour rate into a monthly figure you can defend before you commit.
What this paper shows. viabandwidth lists providers, their models and their operator type, and links to each provider's own on-demand rates where they publish them. It never invents a price or presents a scraped aggregate as its own.

viabandwidth is the verified directory of GPU compute, colocation and carrier networks.

Counts are drawn from the viabandwidth GPU directory, regenerated nightly, where a listing is verified against independent network and infrastructure signals before it is trusted. This paper reports provider counts and the accelerators operators publish. It never quotes a price, because the rate you pay depends on your term, region and workload, so every figure here is a count you can check and never a number we invented.

© 2026 viabandwidth (Steven Higashi). All rights reserved. The figures and charts in this paper may be quoted with clear attribution to viabandwidth.

Cite as: viabandwidth, On-Demand GPU Pricing, viabandwidth.com/gpu/buyers-guide/on-demand-pricing.

viabandwidth.com · [email protected]