On-Demand GPU Pricing
On-demand, spot and committed are a gradient of price against certainty. This is what on-demand really means, how the three modes differ, and when paying on-demand is the right call.
The short answer
On-demand means you pay per unit of time for a GPU instance with no commitment, spin it up when you need it and stop paying when you stop it. That flexibility is the most expensive way to buy a GPU hour and it is the right way for work that is variable or not yet understood. Spot capacity sits below it in price and above it in risk, and committed terms sit below it in price and above it in obligation. Reading a rate correctly means knowing which of the three you are looking at.
01 · The modeWhat on-demand actually means
On-demand pricing is the pay-as-you-go rate for a GPU instance. You start it when you want it, you are billed by the hour or the second while it runs and the meter stops when you stop the machine, with no term to sign and no capacity reserved in advance. That is the model that lets you scale with demand and walk away cleanly.
The convenience carries a premium. Because the provider is holding capacity ready for you to claim at any moment and asking nothing in return, the on-demand rate is the highest of the standard pricing modes. A workload that runs almost all the time on on-demand is quietly paying for a flexibility it never uses, which is the single most common way GPU bills end up larger than they needed to be.
02 · The gradientOn-demand, spot and committed
Think of the three modes as a gradient of price against certainty. Spot is the cheapest, because you are renting hardware the operator would otherwise leave idle and accepting that it can be reclaimed with little notice, which suits interruptible batch work that checkpoints and resumes. Committed terms are also cheaper than on-demand, because you are promising to pay for a fixed slice of capacity over a year or more. On-demand sits in the middle on price and at the top on flexibility, so almost no efficient setup uses one mode for everything.
03 · RegionWhere on-demand capacity sits
The accelerator model is the largest lever on an on-demand rate and region is the next, because capacity is priced differently around the world and a cheaper region far from your users can cost you in latency what it saves in rate. The counts below are the number of tracked providers with indexed capacity in each country.
| Country | Providers indexed |
|---|---|
| United States | 62 |
| Netherlands | 48 |
| Germany | 43 |
| India | 39 |
| United Kingdom | 37 |
| Singapore | 31 |
| Canada | 26 |
| Japan | 23 |
A cheaper on-demand rate in a distant region is only cheaper until latency or egress is counted, so treat geography as a filter on the shortlist rather than a way to chase the lowest number.
What the on-demand catalog actually offers
The accelerator model is the largest single lever on an on-demand rate, and the depth of supply behind each card decides how much that rate can be shopped. The chart below is the number of tracked providers listing each accelerator.
| Accelerator | Providers listing it |
|---|---|
| H100 | 224 |
| A100 | 181 |
| H200 | 147 |
| B200 | 102 |
| L40S | 101 |
| RTX 4090 | 76 |
| V100 | 64 |
| A40 | 49 |
On-demand rates move most on the cards with the deepest supply, so a workload that can run on an A100 or an L40S will find the on-demand premium over spot far easier to swallow than one that needs the scarce frontier silicon.
04 · The callWhen on-demand is the right choice
- 01Use on-demand while the workload is variable or unknown. Prototyping, model evaluation and genuinely spiky traffic are what the premium buys.
- 02Move the steady part off on-demand. Once the shape settles, a committed floor or a dedicated server is cheaper for the predictable work.
- 03Read what the rate includes. Whole-machine flat pricing and accelerator-only pricing with separate storage and egress can diverge sharply on a real month.
- 04Price a month, not an hour. Turn the per-hour rate into a monthly figure you can defend before you commit.
Present this paper
One sponsor can present this buyer's guide on its own, with a logo and a presented-by credit on the paper here and on the downloadable PDF, reaching buyers weighing on-demand against spot and committed GPU capacity.
- ▸Logo and presented-by credit on the paper and the downloadable PDF
- ▸A credit in the Related research module shown beside the GPU operator profiles buyers already read
- ▸Exclusive to a single sponsor per edition, never shared
- ▸Independent by design, so the credit is labelled and never changes the counts or the verification behind them
Or write to [email protected].