viabandwidthGPU Compute
viabandwidthBuyer's Guide

Rent a GPU for LLMs

Inference and training put opposite demands on a machine, so the right rental for one is usually the wrong rental for the other. This is how to size the accelerator to the model and match the machine to the job.

Sponsor this paper →
1,009
GPU providers indexed
582
With published accelerators
19
Accelerator models tracked
144
Network-verified direct operators

The short answer

Serving a model is a latency problem and training one is a throughput problem, so they reward opposite machines. Inference wants a well-sized single card in a region near your users, priced on-demand so you scale with traffic. Training wants raw memory and fast interconnect between cards, priced on a committed term because a serious run is not something you start and stop. Getting the machine to match the job is what separates a bill you can defend from one you cannot.

01 · The jobInference and training want different machines

Serving a model to users rewards a card with enough memory to hold the model and its context, sited close to the people calling it, billed only for the capacity your traffic uses. A single well-chosen accelerator often does the whole job and interruptible capacity is fine as long as you can fail over, so the cheap end of the market is genuinely open to you here.

Training or fine-tuning is a different constraint. Now the questions are how much model and how large a batch you can fit in memory and how fast the cards can talk to each other when the model is too big for one of them. A run that spans several GPUs lives or dies on the interconnect between them, so a training rental is really a rental of a tightly coupled cluster rather than a card that happens to come in a pack.

02 · SizingMatch the accelerator to the model

Settle memory first. A model has to fit along with its activations and the context it works over, so take the parameter count, account for the precision you will run at and leave real headroom rather than the theoretical minimum. Quantising to a lower precision buys room and lets a smaller card carry a bigger model, at a quality cost you should measure rather than assume.

Once memory is settled the family follows the job. An H100 or H200 is the workhorse for demanding training and high-throughput inference, with the H200 carrying more memory for the largest models, while a B200 sits above them for the heaviest work. An A100 stays a capable and often cheaper option and a card like the L40S is a sensible inference choice when you do not need a flagship.

03 · SupplyWhat the market actually offers

Once you know the card, supply decides how much choice you have. The counts below are the number of tracked providers listing each accelerator, which is where competition for your workload is deepest.

H100224A100181H200147B200102L40S101RTX 409076V10064A4049TRACKED PROVIDERS LISTING EACH ACCELERATOR
Number of tracked providers listing each accelerator. H100 and A100 supply runs deepest, so a workload that fits them has the widest choice of operator.
AcceleratorProviders listing it
H100224
A100181
H200147
B200102
L40S101
RTX 409076
V10064
A4049

H100 and A100 remain the workhorses by a wide margin, so a model that fits them has the widest supply and the most leverage on price. The newest silicon spreads fast underneath them, yet the rack-scale systems stay concentrated among fewer operators, where availability rather than price becomes the binding constraint.

Where the H100 supply sits for inference latency

For serving, region decides latency, so the geography of the leading LLM card matters as much as the count. The chart below is the number of operators that state H100 capacity in each country.

India16Germany13United States12Australia10Canada10Netherlands9Singapore8Japan8OPERATORS STATING H100 CAPACITY, BY COUNTRY
Operators stating H100 capacity by country. A serving deployment can be sited near its users without leaving the H100 behind.
CountryOperators with H100
India16
Germany13
United States12
Australia10
Canada10
Netherlands9
Singapore8
Japan8

Depth outside the United States means an inference fleet can sit close to the users calling it, in Europe or India, without dropping to a weaker card. For training, geography matters less than the interconnect quality inside a single cluster.

04 · ProcurementWhat to confirm before you commit

  • 01For inference, confirm the region and the network path to your users. A cheaper card two continents away can lose you more in latency than it saves in rate.
  • 02For training, confirm the interconnect and the storage throughput. A fast cluster starved of data is an expensive way to wait, and not every operator that lists a flagship can deliver a coherent multi-node cluster.
  • 03Price a real month, not a single hour. Ask each operator for the all-in cost of a workload shaped like yours.
  • 04Match the term to your confidence. On-demand while you are still learning the workload, a committed floor once the shape is settled.
What this paper shows. viabandwidth lists provider counts, accelerator models and operator type, and links to each operator's own live rates where they publish them. It never quotes a price, because the number you pay depends on the term, the region and the shape of your workload.

viabandwidth is the verified directory of GPU compute, colocation and carrier networks.

Counts are drawn from the viabandwidth GPU directory, regenerated nightly, where a listing is verified against independent network and infrastructure signals before it is trusted. This paper reports provider counts and the accelerators operators publish. It never quotes a price, because the rate you pay depends on your term, region and workload, so every figure here is a count you can check and never a number we invented.

© 2026 viabandwidth (Steven Higashi). All rights reserved. The figures and charts in this paper may be quoted with clear attribution to viabandwidth.

Cite as: viabandwidth, Rent a GPU for LLMs, viabandwidth.com/gpu/buyers-guide/rent-gpu-llm.

viabandwidth.com · [email protected]