NVIDIA B200 (Blackwell)
NVIDIA Blackwell · dual-die · TSMC 4NP
VRAM192 GB HBM3e (~180 GB usable) Bandwidth~8 TB/s TDP1000 W (up to 1200 W) — liquid cooling for most builds Interconnect5th-gen NVLink 1.8 TB/s; PCIe Gen5
Best forNew 8–64 GPU training clusters and high-throughput LLM inference; 70B at FP16 or up to ~120B at FP8 on a single GPU.
Second-generation Transformer Engine with native FP4 gives roughly 4× the inference of an H100 and ~2–2.5× training. Supply is the constraint — large order backlogs through mid-2026 mean cloud rental is the fast path, and a US export-control variant exists.
NVIDIA B300 (Blackwell Ultra)
NVIDIA Blackwell Ultra
VRAM288 GB HBM3e Bandwidth~8 TB/s TDP~1400 W — liquid cooling Interconnect5th-gen NVLink; ConnectX-8 (doubled inter-node)
Best forWhen a single GPU must hold a 70B+ model without quantisation, or FP4 throughput is the hard constraint.
A higher-binned B200 with ~50% more VRAM and ~67% more FP4 compute; worth it only when the extra memory or networking is the actual bottleneck.
NVIDIA H200 (Hopper)
NVIDIA Hopper (memory upgrade on H100)
VRAM141 GB HBM3e Bandwidth~4.8 TB/s TDP700 W — air-coolable Interconnect4th-gen NVLink; PCIe Gen5
Best forThe pragmatic default: ample memory, mature software, lower power, available now.
Handles the large majority of training and inference today; the right pick when 141 GB is enough and you do not need Blackwell's FP4 throughput.
NVIDIA H100 (Hopper)
NVIDIA Hopper
VRAM80 GB HBM3 Bandwidth~3.35 TB/s TDP700 W — air-coolable Interconnect4th-gen NVLink; PCIe Gen5
Best forCost-effective workhorse for fine-tuning and inference at sensible model sizes.
Still capable and abundant; prices keep falling as Blackwell supply grows, which makes it strong value when 80 GB fits the job.
AMD Instinct MI355X (CDNA 4)
AMD CDNA 4 · chiplet · TSMC 3nm
VRAM288 GB HBM3e Bandwidth~8 TB/s TDP~1400 W — liquid cooling (MI350X air-cooled, lower TBP) InterconnectInfinity Fabric mesh (~1.075 TB/s peer); ROCm 7 software
Best forMemory-bound inference and models that benefit from fitting on one GPU (500B+ parameters in 288 GB).
Competitive with B200 on single-node FP4 inference and near-parity on Llama-class fine-tuning; the trade-offs are the smaller ROCm ecosystem and a tighter rental market than NVIDIA's.
AMD Instinct MI300X (CDNA 3)
AMD CDNA 3 · chiplet
VRAM192 GB HBM3 Bandwidth~5.3 TB/s TDP750 W InterconnectInfinity Fabric; ROCm
Best forMemory-bound models that do not fit a single H200, where ROCm support is in place.
More memory than an H200 and strong value for the right inference workloads; multi-node scaling and ecosystem maturity trail NVIDIA.