viabandwidthGPU Compute
GPU economics · interactive

The AI Compute Cost Calculator

Compare accelerators by the work they finish, not the hourly sticker price. Enter your workload and the tool returns cost per million tokens, tokens per watt, the number of GPUs your model actually needs and the option that fits your priority.

Illustrative defaults, not sourced yet. Every price, throughput and power figure below is an editable placeholder for a reference ~70B model at FP8. Replace them with your own live quotes and measured benchmarks. This calculator asserts only the arithmetic. The market numbers are yours to supply.

Your numbers are saved on this device.
Layout

What are you planning to run?

Choose the closest example, then adjust any value that does not match.

Start here

Describe the workload

These values update the estimate as you type.

This workload produces about 252.0M output tokens each month and needs an estimated 112 GB of memory.

Recommended fitBased on these inputs
MI300X · 192GB$8.29 / 1M out

Lowest total monthly cost under your demand. It needs 1 accelerator (limited by balanced) at an effective $8.29 per million output tokens.

Monthly cost
$2,088
GPUs needed
1 · balanced
Tokens / watt
2.00
Utilization
3%
Effective cost per 1M output tokens
At your demand, including any idle capacity from rounding up to whole GPUs. Lower is better.
MI300X · 192GB
$8.29
H200 · 141GB
$9.14
B200
$13.14
H100 · 80GB
$14.29
B300
$14.57

Compare and adjust the assumptions

Type over any price, throughput or power figure with your own numbers. Everything recalculates instantly.
Accelerator$/hrTokens/sPower WHBM GBGPUs$/1M out$/1M inTok/WMonthly
H100 · 80GB
Hopper
2$14.29$0.062.08$3,600
H200 · 141GB
Hopper, more HBM
1$9.14$0.062.46$2,304
B200
Blackwell
1$13.14$0.053.06$3,312
B300
Blackwell Ultra
1$14.57$0.042.95$3,672
MI300X · 192GB
AMD
1$8.29$0.062.00$2,088
Monthly output tokens
252.0M
Monthly input tokens
432.0M
Memory to fit model
112 GB
Confidence · illustrative

Results are estimates, not procurement guarantees. They depend on the numbers you enter. Actual performance depends on model architecture, software, precision, batching, utilization, system configuration, networking, cooling and provider terms. Memory-fit and cost figures are exact arithmetic on your inputs. Throughput and price are assumptions you should validate against your own production workload and current quotes. Input-token cost uses your prefill-speed assumption and is reported separately from output-token cost, because output generation is usually the heavier and more expensive step.

Browse verified GPU providers →Compare specific operators

What the AI compute cost calculator does

This free calculator compares AI accelerators by the work they finish rather than the hourly sticker price. You enter your model, your monthly demand and your latency priority, then it returns cost per million tokens, tokens per watt, tokens per dollar, the number of GPUs your model actually needs and an estimated monthly cost. It covers the NVIDIA H100, H200, B200 and B300 alongside the AMD MI300X, and every price and benchmark is an editable input so the answer reflects your own quotes.

How to compare GPUs by cost per token

Cost per million tokens is the hourly infrastructure cost divided by the tokens generated in that hour, multiplied by one million. A GPU that rents for more per hour can still produce tokens more cheaply when it finishes enough additional work in the same time. Memory matters as well, because a model that does not fit on one card needs several of them, and the number of GPUs is usually the largest line on the bill. The calculator handles both effects, so you can see the point where a higher hourly rate becomes the lower monthly cost.

Accelerators you can compare

  • NVIDIA H100 80 GB, the mature baseline for inference and training
  • NVIDIA H200 141 GB, more memory for long context and larger models
  • NVIDIA B200 and B300, Blackwell throughput for high-volume inference
  • AMD MI300X 192 GB, a high-memory alternative

Frequently asked questions

Is the cheapest GPU-hour always the cheapest AI compute?

No. A GPU that rents for more per hour can cost less per token when it finishes more work in that hour, so the honest unit of cost is the price per million tokens rather than the hourly rate.

How do I calculate cost per million tokens?

Divide the hourly infrastructure cost by the tokens generated in that hour, then multiply by one million. Input and output tokens are best kept separate because output generation is usually the heavier step.

How many GPUs does a 70B model need?

A 70B model at FP8 precision needs roughly 112 GB once weights and working memory are counted, so it fits on one 192 GB card and needs two 80 GB cards. Memory fit often sets the GPU count before throughput does.

What is tokens per watt?

Tokens per watt is output tokens per second divided by average system watts. It becomes the deciding metric when a facility runs out of power before it runs out of budget.

Which is cheaper, H100 or B300?

It depends on demand. At low demand the H100 can win on total cost while at high utilization the B300 usually wins on cost per token, because it finishes far more work in the same hour.

Read the full paper →GPU provider directory →