GPU Benchmark Comparator

Compare cloud/training GPUs — H100, H200, A100, B200, GB200, and AMD MI300X — across MLPerf benchmarks, memory bandwidth, TFLOPS, VRAM, TDP, street price, and cloud hourly cost.

Workload preset

Spec comparison

GPU Arch VRAM HBM BW FP16/BF16 FP8 FP4 FP64 TDP Street price Cloud/hr Best for
NVIDIA A100 SXM Ampere 80 GB 2 TB/s 312 T 9.7 T 400 W $12,000 $2/hr Legacy HPC, Training smaller models, Budget cloud spot
NVIDIA H100 SXM Hopper 80 GB 3.35 TB/s 989 T 1,979 T 34 T 700 W $25,000 $3.5/hr Production training, FP8 inference, Mature CUDA ecosystem
NVIDIA H200 SXM Hopper 141 GB 4.8 TB/s 989 T 1,979 T 34 T 700 W $32,000 $5.9/hr Memory-bound inference, 70B+ models, Long context decode
NVIDIA B200 SXM Blackwell 180 GB 7.7 TB/s 2,250 T 4,500 T 9,000 T 37 T 1000 W $40,000 $5.34/hr Next-gen training, FP4 inference, Large-scale clusters
NVIDIA GB200 NVL (per GPU die) Grace Blackwell 186 GB 8 TB/s 2,500 T 5,000 T 10,000 T 37 T 1200 W $55,000 Rack-only Rack-scale reasoning models, CPU-GPU unified memory, NVL72 clusters
AMD Instinct MI300X CDNA 3 192 GB 5.3 TB/s 1,307 T 2,615 T 82 T 750 W $20,000 $2.5/hr Massive single-GPU memory, FP64 HPC, Budget 70B+ inference

All figures are peak/dense accelerator specs and representative benchmark speedups for direct comparison. Real throughput depends on framework, model size, precision, sparsity, and interconnect. Check MLPerf submissions for apples-to-apples results.

MLPerf Training GPT-3 speedup vs A100

NVIDIA A100 SXM 1.0x
NVIDIA H100 SXM 2.6x
NVIDIA H200 SXM 2.9x
NVIDIA B200 SXM 5.5x
NVIDIA GB200 NVL (per GPU die) 6.2x
AMD Instinct MI300X 2.2x

MLPerf Inference Llama 2 70B speedup vs A100

NVIDIA A100 SXM 1.0x
NVIDIA H100 SXM 2.2x
NVIDIA H200 SXM 3.0x
NVIDIA B200 SXM 8.5x
NVIDIA GB200 NVL (per GPU die) 10.0x
AMD Instinct MI300X 2.5x

Cost-per-speedup

Street price divided by MLPerf inference speedup vs A100. Lower is better raw hardware value for inference.

GPU Street price Inference speedup Price per x speedup Cloud $ per x speedup/hr
NVIDIA A100 SXM $12,000 1.0x $12,000 $2.00
NVIDIA H100 SXM $25,000 2.2x $11,364 $1.59
NVIDIA H200 SXM $32,000 3.0x $10,667 $1.97
NVIDIA B200 SXM $40,000 8.5x $4,706 $0.63
NVIDIA GB200 NVL (per GPU die) $55,000 10.0x $5,500
AMD Instinct MI300X $20,000 2.5x $8,000 $1.00

Cloud providers

  • CoreWeave

    H100, H200, B200, GB200 — GPU-first cloud; strong NVLink clusters and spot pricing.

  • Lambda

    A100, H100, H200, B200 — Research-friendly persistent instances and 1-click Jupyter.

  • Vast.ai / RunPod

    A100, H100, H200, MI300X — Consumer and datacenter marketplace; highly variable pricing.

  • Crusoe / DigitalOcean

    H100, MI300X — Sustainable compute; DigitalOcean hosts bare-metal MI300X.

  • Oracle Cloud

    A100, H100, H200, MI300X — Enterprise bare-metal clusters with MI300X and H100 options.

Which GPU should you pick?

LLM pretraining (GPT-3 class)

FP8 throughput and NVLink scaling dominate time-to-convergence.

Top: NVIDIA B200 / GB200 Budget: NVIDIA H100 cluster

70B model inference (batch=1)

Decode is memory-bandwidth-bound; extra VRAM and bandwidth pay off most.

Top: NVIDIA H200 Budget: AMD MI300X

FP4 production serving

Blackwell is the only generation with native FP4 tensor cores.

Top: NVIDIA B200 Budget: NVIDIA H100 (FP8)

FP64 HPC / simulation

MI300X leads FP64 throughput among these options.

Top: AMD MI300X Budget: NVIDIA A100

Rack-scale reasoning cluster

NVL72 gives single-domain scaling and unified memory ideal for reasoning.

Top: NVIDIA GB200 NVL72 Budget: NVIDIA H200 cluster

Caveats

  • Peak TFLOPS are not real-world throughput. Framework, model size, batch size, quantization, sparsity, and interconnect all change results.
  • Cloud hourly rates are directional and can swing 2-5x based on spot, contract, and availability.
  • MLPerf speedups are approximate normalized comparisons vs A100; consult the latest submission tables for exact configs.
  • GB200 pricing is typically rack/reservation; street price shown is a per-GPU-die estimate for TCO modeling.

Last updated: 2026-08-07.

Frequently asked questions

Which GPU is best for LLM training?

For large-scale training, NVIDIA B200 or GB200 lead on FP8/FP4 throughput and NVLink bandwidth. If budget or availability is tight, H100 clusters remain the proven choice. AMD MI300X can train but typically trails H100 in real-world frameworks unless you optimize ROCm carefully.

Is H200 worth the premium over H100?

Yes if you serve 70B+ parameter models, run long-context inference, or are memory-bandwidth-bound. The compute silicon is identical; you are paying for 141 GB HBM3e at 4.8 TB/s. For compute-bound training that fits in 80 GB, H100 is the better value.

When does B200 beat H200?

B200 pulls ahead when FP4 is available in your model and serving stack, and when you need >141 GB VRAM or >4.8 TB/s bandwidth. MLPerf shows B200 roughly 2-4x H100/H200 on supported inference and ~2x on training; real margins depend on framework maturity.

What is GB200 vs B200?

B200 is a standalone Blackwell GPU that slots into HGX servers. GB200 is the Grace Blackwell Superchip (two B200 dies + one Grace CPU on a package), primarily sold in NVL72 rack systems for CPU-GPU unified memory and rack-scale inference.

Should I consider AMD MI300X?

Consider MI300X when memory capacity per GPU is the bottleneck (192 GB) or when FP64 HPC throughput matters. Real-world AI throughput is often 75-90% of equivalent NVIDIA solutions due to software maturity, so factor engineering time into TCO.

🚀 Get AI automation insights daily

15:00 MST. One-click unsubscribe.

Subscribe