NVIDIA’s Blackwell Ultra is no longer a roadmap slide — it’s shipping, and the benchmark numbers are in. The GB300 NVL72 just set records across every MLPerf Inference v5.1 category. For anyone building or buying AI infrastructure, this matters now.
What Happened
Announced at GTC in March 2025 and now deployed by hyperscalers, the Blackwell Ultra architecture (GB300/B300) is the successor to the original Blackwell (GB200/B200).
Key specs vs its predecessor:
| GB200 (Blackwell) | GB300 (Blackwell Ultra) | |
|---|---|---|
| HBM3e memory per GPU | 192 GB | 288 GB |
| NVFP4 compute | 10 PFLOPS | 15 PFLOPS (+50%) |
| PCIe | Gen 5 (64 GB/s) | Gen 6 (128 GB/s) |
| GPU TDP | 1,200W | 1,400W |
| NVL72 rack inference | baseline | +45% DeepSeek-R1 throughput |
The GB300 NVL72 rack packs 72 Blackwell Ultra GPUs and 36 Grace CPUs, acting as a single 1.1 exaFLOPS compute unit. AWS, Google Cloud, Azure, and Oracle are all among the first to offer Blackwell Ultra instances.
In MLPerf Inference v5.1, the GB300 NVL72 delivered 45% higher DeepSeek-R1 offline throughput vs the previous GB200 NVL72 — and set records on every new benchmark added to the suite, including Llama 3.1 405B and Whisper.
Read the full MLPerf results →
My Take
The 45% MLPerf jump is significant — but the more interesting number is 50x revenue opportunity vs Hopper. NVIDIA is pitching Blackwell Ultra not just as faster hardware but as a business model shift for cloud providers: charge premium rates for time-sensitive inference (reasoning models, agentic workflows) and extract far more revenue per rack.
The NVFP4 format is worth watching. It’s proprietary — not standard IEEE FP4 — which means applications running on Blackwell Ultra are increasingly tied to NVIDIA’s ecosystem. That’s a deliberate choice. Better accuracy than INT4, close to BF16, but you’re fully on the NVIDIA stack.
For developers: if you’re running DeepSeek-R1, Llama 3.1 405B, or any reasoning model at scale, the throughput and cost-per-token story is now meaningfully better than Hopper. The question is whether your cloud provider has GB300 instances available yet.
What to Watch
- Vera Rubin (2026): Next architecture after Blackwell Ultra — 50 PFLOPS inference per package, NVL144 rack. Already announced at GTC.
- NVFP4 ecosystem adoption: More frameworks need to support this natively. NVIDIA’s TensorRT-LLM does — third-party stacks lag.
- Cloud pricing: Providers haven’t published GB300 rates publicly. Expect premium. Real test is TCO vs GB200 at scale.
