AI Chips · Product page

NVIDIA H200 (Hopper): The Complete Guide

Same compute die as H100, a bigger and faster memory system bolted on. Simple to expert, every number links to its source.

Ask an AI about this chip:CAGPX
NVIDIA H200 10 sections · 2 levels 4 linked sources Checked September 2026
How to read this page. Each section starts Simple, then goes to a Deep dive: stop wherever you have what you need. The small numbers are sources: click one to open the original document. Where the maker has not published a figure -- a price, a die size, a factory address -- this page says so rather than estimate it.

1.At a glance

SimpleStart here

The H200 is NVIDIA’s upgrade to H100 announced 13 November 2023 1. It uses the identical compute die as H100 — the change is entirely in the memory system: nearly double the capacity, close to 45% more bandwidth.

Deep diveThe technical detail

Codename GH100 (same die family as H100), architecture Hopper. General availability followed in Q2 2024 from system manufacturers and cloud providers 1. Ships as SXM, the air-cooled H200 NVL PCIe variant, and inside the GH200 Grace Hopper Superchip 1.

2.Launch and history

SimpleStart here

H200 exists because HBM memory capacity and bandwidth, not raw compute, had become the bottleneck for running the largest language models. NVIDIA’s own release ties the launch directly to inference speed on a named model.

Deep diveThe technical detail

NVIDIA’s own press release states H200 delivers “nearly doubling inference speed on Llama 2, a 70 billion-parameter LLM,” compared with H100 1 — a company-published, model-specific performance claim, not an independent benchmark. The release lists ASRock Rack, ASUS, Dell, Eviden, GIGABYTE, HPE, Ingrasys, Lenovo, QCT, Supermicro, Wistron and Wiwynn as system partners, and AWS, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure, CoreWeave, Lambda and Vultr as cloud partners 1.

3.What’s inside it

SimpleStart here

H200 carries over every Hopper architecture feature from H100 — the Transformer Engine, FP8 support, confidential computing — unchanged. Nothing about the compute logic is new; only the memory attached to it changed.

Deep diveThe technical detail

H200’s architectural features are the Hopper features already documented for H100: Transformer Engine, fourth-generation NVLink, second-generation Multi-Instance GPU, DPX instructions for dynamic programming, and built-in confidential computing 2. NVIDIA’s H200 materials do not describe any new compute-side architectural feature — the announcement is framed entirely around the memory upgrade 1.

4.Full spec table

SimpleStart here

The compute specs are identical to H100. What changed: 141GB of HBM3e memory (up from 80GB) at 4.8 TB/s bandwidth (up from H100’s 3.35 TB/s) 3.

Deep diveThe technical detail
SpecH200 SXM
Process nodeTSMC 4N (shared GH100 die with H100) 2
Transistors80 billion (shared die) 2
Die size814 mm² (shared die) 2
Memory141GB HBM3e (was 80GB HBM3 on H100)
Memory bandwidth4.8 TB/s (was 3.35 TB/s on H100)
TDPUp to 700 W (SXM)
FP64 / FP64 Tensor34 / 67 TFLOPS
FP3267 TFLOPS
TF32 Tensor (sparse)989 TFLOPS
BF16/FP16 Tensor (sparse)1,979 TFLOPS
FP8 / INT8 Tensor (sparse)3,958 TFLOPS / TOPS
NVLink900 GB/s
PCIeGen5, 128 GB/s
Form factorsSXM; H200 NVL (600W air-cooled PCIe); inside the GH200 Grace Hopper Superchip

Figures per NVIDIA’s own H200 product page 3. Every Tensor Core FLOPS figure here (FP64, FP64 Tensor, FP32, TF32, BF16/FP16, FP8, INT8) is numerically identical to H100 SXM5’s published figures — see the “Hidden in plain sight” section.

5.Where it’s made

SimpleStart here

H200 uses the same physical die as H100, made by TSMC on the custom 4N process 2. Only the memory stacked next to the die changed.

Deep diveThe technical detail

Because H200 is a memory-subsystem revision of the H100 GH100 die, its foundry and process are identical to H100’s: TSMC, custom 4N 2. Neither NVIDIA nor TSMC publishes individual fab locations for either chip.

6.Which systems use it

SimpleStart here

HGX H200 server boards (4- and 8-GPU) and the GH200 Grace Hopper Superchip with H200-class memory 1. Offered by AWS, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure, CoreWeave, Lambda and Vultr, per NVIDIA’s own release.

Deep diveThe technical detail

NVIDIA’s own release names its system partners as ASRock Rack, ASUS, Dell, Eviden, GIGABYTE, HPE, Ingrasys, Lenovo, QCT, Supermicro, Wistron and Wiwynn, and its cloud partners as AWS, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure, CoreWeave, Lambda and Vultr 1. The GH200 Grace Hopper Superchip pairs a Grace CPU with H200-class HBM3e memory in a single package 1.

7.Official pricing

SimpleStart here

NVIDIA has not published a unit or system price for H200. CoreWeave, an official NVIDIA cloud partner, lists HGX H200 (8 GPUs) at $50.44/hr on-demand and $20.93/hr spot, or $6.31/hr for a single-GPU inference instance 4.

Deep diveThe technical detail

This is one of the few chips on this site where a cloud partner’s own published rate card gives a specific, checkable number rather than a general “contact sales” page — CoreWeave’s pricing page states these figures directly 4. AWS, Azure and Google Cloud also offer H200 instances but their exact rates were not independently re-verified for this page; check each provider’s own pricing page for current figures, since cloud rates change.

8.Real-world performance

SimpleStart here

NVIDIA’s own claim: “nearly doubling inference speed on Llama 2, a 70 billion-parameter LLM” compared with H100 1 — a company-stated, single-model figure, not a general multiplier and not an independent benchmark.

Deep diveThe technical detail

Because H200’s compute Tensor-core throughput figures are identical to H100’s, its real-world performance gains come almost entirely from serving larger models and longer contexts without running out of memory, and from the higher memory bandwidth feeding the unchanged compute units faster. NVIDIA’s Llama-2-70B inference claim 1 is consistent with a memory-bandwidth-bound workload benefiting from the 43% bandwidth increase (3.35 TB/s to 4.8 TB/s).

9.What came before, what came next

SimpleStart here

Came before: H100, same compute die. Came after: B200 (Blackwell architecture), announced March 2024.

Deep diveThe technical detail

H200 is not a new architecture generation — it is a memory-upgraded H100. The next architectural generation NVIDIA announced is Blackwell, with the B200 GPU revealed at GTC in March 2024 1 (context from the same period; see the B200 page for its own primary sources).

10.Hidden in plain sight

SimpleStart here

H200 is compute-identical to H100. Every single Tensor Core FLOPS figure NVIDIA publishes for the two chips — FP64, FP64 Tensor, FP32, TF32, BF16/FP16, FP8, INT8 — is exactly the same number. The entire “H200” product is a memory-subsystem swap on the same 80-billion-transistor, 814mm² die.

Deep diveThe technical detail

Comparing NVIDIA’s own H100 and H200 specification pages side by side 23: FP64 34/67 TFLOPS, FP32 67 TFLOPS, TF32 989 TFLOPS (sparse), BF16/FP16 1,979 TFLOPS (sparse), FP8/INT8 3,958 TFLOPS/TOPS (sparse), and 900 GB/s NVLink are identical across both chips. NVIDIA did not change the compute die at all between H100 and H200 — it changed only the memory type (HBM3 to HBM3e), capacity (80GB to 141GB) and bandwidth (3.35 TB/s to 4.8 TB/s).

11.Sources

4 sources, checked September 2026. Where NVIDIA or another maker has not published a figure, this page says so rather than estimate it.

  1. NVIDIA Investor Relations: NVIDIA Supercharges Hopper, the World’s Leading AI Computing PlatformOfficial
  2. NVIDIA: NVIDIA H100 Tensor Core GPU Architecture WhitepaperOfficial
  3. NVIDIA: H200 Tensor Core GPU product pageOfficial
  4. CoreWeave: Cloud GPU PricingOfficial