GPU Benchmark Comparator
Compare cloud/training GPUs — H100, H200, A100, B200, GB200, and AMD MI300X — across MLPerf benchmarks, memory bandwidth, TFLOPS, VRAM, TDP, street price, and cloud hourly cost.
Workload preset
Spec comparison
| GPU | Arch | VRAM | HBM BW | FP16/BF16 | FP8 | FP4 | FP64 | TDP | Street price | Cloud/hr | Best for |
|---|---|---|---|---|---|---|---|---|---|---|---|
| NVIDIA A100 SXM | Ampere | 80 GB | 2 TB/s | 312 T | — | — | 9.7 T | 400 W | $12,000 | $2/hr | Legacy HPC, Training smaller models, Budget cloud spot |
| NVIDIA H100 SXM | Hopper | 80 GB | 3.35 TB/s | 989 T | 1,979 T | — | 34 T | 700 W | $25,000 | $3.5/hr | Production training, FP8 inference, Mature CUDA ecosystem |
| NVIDIA H200 SXM | Hopper | 141 GB | 4.8 TB/s | 989 T | 1,979 T | — | 34 T | 700 W | $32,000 | $5.9/hr | Memory-bound inference, 70B+ models, Long context decode |
| NVIDIA B200 SXM | Blackwell | 180 GB | 7.7 TB/s | 2,250 T | 4,500 T | 9,000 T | 37 T | 1000 W | $40,000 | $5.34/hr | Next-gen training, FP4 inference, Large-scale clusters |
| NVIDIA GB200 NVL (per GPU die) | Grace Blackwell | 186 GB | 8 TB/s | 2,500 T | 5,000 T | 10,000 T | 37 T | 1200 W | $55,000 | Rack-only | Rack-scale reasoning models, CPU-GPU unified memory, NVL72 clusters |
| AMD Instinct MI300X | CDNA 3 | 192 GB | 5.3 TB/s | 1,307 T | 2,615 T | — | 82 T | 750 W | $20,000 | $2.5/hr | Massive single-GPU memory, FP64 HPC, Budget 70B+ inference |
All figures are peak/dense accelerator specs and representative benchmark speedups for direct comparison. Real throughput depends on framework, model size, precision, sparsity, and interconnect. Check MLPerf submissions for apples-to-apples results.
MLPerf Training GPT-3 speedup vs A100
MLPerf Inference Llama 2 70B speedup vs A100
Cost-per-speedup
Street price divided by MLPerf inference speedup vs A100. Lower is better raw hardware value for inference.
| GPU | Street price | Inference speedup | Price per x speedup | Cloud $ per x speedup/hr |
|---|---|---|---|---|
| NVIDIA A100 SXM | $12,000 | 1.0x | $12,000 | $2.00 |
| NVIDIA H100 SXM | $25,000 | 2.2x | $11,364 | $1.59 |
| NVIDIA H200 SXM | $32,000 | 3.0x | $10,667 | $1.97 |
| NVIDIA B200 SXM | $40,000 | 8.5x | $4,706 | $0.63 |
| NVIDIA GB200 NVL (per GPU die) | $55,000 | 10.0x | $5,500 | — |
| AMD Instinct MI300X | $20,000 | 2.5x | $8,000 | $1.00 |
Cloud providers
- CoreWeave
H100, H200, B200, GB200 — GPU-first cloud; strong NVLink clusters and spot pricing.
- Lambda
A100, H100, H200, B200 — Research-friendly persistent instances and 1-click Jupyter.
- Vast.ai / RunPod
A100, H100, H200, MI300X — Consumer and datacenter marketplace; highly variable pricing.
- Crusoe / DigitalOcean
H100, MI300X — Sustainable compute; DigitalOcean hosts bare-metal MI300X.
- Oracle Cloud
A100, H100, H200, MI300X — Enterprise bare-metal clusters with MI300X and H100 options.
Which GPU should you pick?
LLM pretraining (GPT-3 class)
FP8 throughput and NVLink scaling dominate time-to-convergence.
70B model inference (batch=1)
Decode is memory-bandwidth-bound; extra VRAM and bandwidth pay off most.
FP4 production serving
Blackwell is the only generation with native FP4 tensor cores.
FP64 HPC / simulation
MI300X leads FP64 throughput among these options.
Rack-scale reasoning cluster
NVL72 gives single-domain scaling and unified memory ideal for reasoning.
Caveats
- Peak TFLOPS are not real-world throughput. Framework, model size, batch size, quantization, sparsity, and interconnect all change results.
- Cloud hourly rates are directional and can swing 2-5x based on spot, contract, and availability.
- MLPerf speedups are approximate normalized comparisons vs A100; consult the latest submission tables for exact configs.
- GB200 pricing is typically rack/reservation; street price shown is a per-GPU-die estimate for TCO modeling.
Last updated: 2026-08-07.
Frequently asked questions
Which GPU is best for LLM training?▼
For large-scale training, NVIDIA B200 or GB200 lead on FP8/FP4 throughput and NVLink bandwidth. If budget or availability is tight, H100 clusters remain the proven choice. AMD MI300X can train but typically trails H100 in real-world frameworks unless you optimize ROCm carefully.
Is H200 worth the premium over H100?▼
Yes if you serve 70B+ parameter models, run long-context inference, or are memory-bandwidth-bound. The compute silicon is identical; you are paying for 141 GB HBM3e at 4.8 TB/s. For compute-bound training that fits in 80 GB, H100 is the better value.
When does B200 beat H200?▼
B200 pulls ahead when FP4 is available in your model and serving stack, and when you need >141 GB VRAM or >4.8 TB/s bandwidth. MLPerf shows B200 roughly 2-4x H100/H200 on supported inference and ~2x on training; real margins depend on framework maturity.
What is GB200 vs B200?▼
B200 is a standalone Blackwell GPU that slots into HGX servers. GB200 is the Grace Blackwell Superchip (two B200 dies + one Grace CPU on a package), primarily sold in NVL72 rack systems for CPU-GPU unified memory and rack-scale inference.
Should I consider AMD MI300X?▼
Consider MI300X when memory capacity per GPU is the bottleneck (192 GB) or when FP64 HPC throughput matters. Real-world AI throughput is often 75-90% of equivalent NVIDIA solutions due to software maturity, so factor engineering time into TCO.