The AI Compute Cost Calculator
Compare accelerators by the work they finish, not the hourly sticker price. Enter your workload and the tool returns cost per million tokens, tokens per watt, the number of GPUs your model actually needs and the option that fits your priority.
Illustrative defaults, not sourced yet. Every price, throughput and power figure below is an editable placeholder for a reference ~70B model at FP8. Replace them with your own live quotes and measured benchmarks. This calculator asserts only the arithmetic. The market numbers are yours to supply.
What are you planning to run?
Choose the closest example, then adjust any value that does not match.
Describe the workload
These values update the estimate as you type.
This workload produces about 252.0M output tokens each month and needs an estimated 112 GB of memory.
Lowest total monthly cost under your demand. It needs 1 accelerator (limited by balanced) at an effective $8.29 per million output tokens.
Results are estimates, not procurement guarantees. They depend on the numbers you enter. Actual performance depends on model architecture, software, precision, batching, utilization, system configuration, networking, cooling and provider terms. Memory-fit and cost figures are exact arithmetic on your inputs. Throughput and price are assumptions you should validate against your own production workload and current quotes. Input-token cost uses your prefill-speed assumption and is reported separately from output-token cost, because output generation is usually the heavier and more expensive step.
What the AI compute cost calculator does
This free calculator compares AI accelerators by the work they finish rather than the hourly sticker price. You enter your model, your monthly demand and your latency priority, then it returns cost per million tokens, tokens per watt, tokens per dollar, the number of GPUs your model actually needs and an estimated monthly cost. It covers the NVIDIA H100, H200, B200 and B300 alongside the AMD MI300X, and every price and benchmark is an editable input so the answer reflects your own quotes.
How to compare GPUs by cost per token
Cost per million tokens is the hourly infrastructure cost divided by the tokens generated in that hour, multiplied by one million. A GPU that rents for more per hour can still produce tokens more cheaply when it finishes enough additional work in the same time. Memory matters as well, because a model that does not fit on one card needs several of them, and the number of GPUs is usually the largest line on the bill. The calculator handles both effects, so you can see the point where a higher hourly rate becomes the lower monthly cost.
Accelerators you can compare
- NVIDIA H100 80 GB, the mature baseline for inference and training
- NVIDIA H200 141 GB, more memory for long context and larger models
- NVIDIA B200 and B300, Blackwell throughput for high-volume inference
- AMD MI300X 192 GB, a high-memory alternative
Frequently asked questions
Is the cheapest GPU-hour always the cheapest AI compute?
No. A GPU that rents for more per hour can cost less per token when it finishes more work in that hour, so the honest unit of cost is the price per million tokens rather than the hourly rate.
How do I calculate cost per million tokens?
Divide the hourly infrastructure cost by the tokens generated in that hour, then multiply by one million. Input and output tokens are best kept separate because output generation is usually the heavier step.
How many GPUs does a 70B model need?
A 70B model at FP8 precision needs roughly 112 GB once weights and working memory are counted, so it fits on one 192 GB card and needs two 80 GB cards. Memory fit often sets the GPU count before throughput does.
What is tokens per watt?
Tokens per watt is output tokens per second divided by average system watts. It becomes the deciding metric when a facility runs out of power before it runs out of budget.
Which is cheaper, H100 or B300?
It depends on demand. At low demand the H100 can win on total cost while at high utilization the B300 usually wins on cost per token, because it finishes far more work in the same hour.