2.Launch and history
SimpleStart here
H200 exists because HBM memory capacity and bandwidth, not raw compute, had become the bottleneck for running the largest language models. NVIDIA’s own release ties the launch directly to inference speed on a named model.
Deep diveThe technical detail
NVIDIA’s own press release states H200 delivers “nearly doubling inference speed on Llama 2, a 70 billion-parameter LLM,” compared with H100 1 — a company-published, model-specific performance claim, not an independent benchmark. The release lists ASRock Rack, ASUS, Dell, Eviden, GIGABYTE, HPE, Ingrasys, Lenovo, QCT, Supermicro, Wistron and Wiwynn as system partners, and AWS, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure, CoreWeave, Lambda and Vultr as cloud partners 1.
6.Which systems use it
SimpleStart here
HGX H200 server boards (4- and 8-GPU) and the GH200 Grace Hopper Superchip with H200-class memory 1. Offered by AWS, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure, CoreWeave, Lambda and Vultr, per NVIDIA’s own release.
Deep diveThe technical detail
NVIDIA’s own release names its system partners as ASRock Rack, ASUS, Dell, Eviden, GIGABYTE, HPE, Ingrasys, Lenovo, QCT, Supermicro, Wistron and Wiwynn, and its cloud partners as AWS, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure, CoreWeave, Lambda and Vultr 1. The GH200 Grace Hopper Superchip pairs a Grace CPU with H200-class HBM3e memory in a single package 1.
10.Hidden in plain sight
SimpleStart here
H200 is compute-identical to H100. Every single Tensor Core FLOPS figure NVIDIA publishes for the two chips — FP64, FP64 Tensor, FP32, TF32, BF16/FP16, FP8, INT8 — is exactly the same number. The entire “H200” product is a memory-subsystem swap on the same 80-billion-transistor, 814mm² die.
Deep diveThe technical detail
Comparing NVIDIA’s own H100 and H200 specification pages side by side 23: FP64 34/67 TFLOPS, FP32 67 TFLOPS, TF32 989 TFLOPS (sparse), BF16/FP16 1,979 TFLOPS (sparse), FP8/INT8 3,958 TFLOPS/TOPS (sparse), and 900 GB/s NVLink are identical across both chips. NVIDIA did not change the compute die at all between H100 and H200 — it changed only the memory type (HBM3 to HBM3e), capacity (80GB to 141GB) and bandwidth (3.35 TB/s to 4.8 TB/s).