How to read this page. Each section starts Simple, then goes to a Deep dive: stop wherever you have what you need. The small numbers are sources: click one to open the original document. Where the maker has not published a figure -- a price, a die size, a factory address -- this page says so rather than estimate it.
1.At a glance
SimpleStart here
The GB200 is not one chip — it is a “Superchip”: one NVIDIA Grace CPU plus two Blackwell GPUs on a single board, announced 18 March 2024 alongside B200 1. NVIDIA sells it both as an individual Superchip and scaled into a 72-GPU rack called the GB200 NVL72.
Deep diveThe technical detail
Design: 1× Grace (Arm CPU) + 2× Blackwell GPU, connected by NVLink-C2C. The NVL72 rack combines 36 Grace CPUs and 72 Blackwell GPUs into one NVLink domain 2. This is NVIDIA’s first rack-scale product marketed as behaving like a single accelerator rather than a cluster of servers.
2.Launch and history
SimpleStart here
GB200 exists because a rack full of separate GPU servers, however fast individually, hits a wall when a model is too large to fit in one server’s memory and has to be split across many machines with slower networking between them. NVIDIA’s answer was to fuse CPU, GPU and rack-scale networking into one design from the start.
Deep diveThe technical detail
Announced at GTC on 18 March 2024, alongside B200 12. At the same keynote, describing the NVL72 rack, Jensen Huang said: “one giant GPU” — “a multi-node, liquid-cooled, rack-scale system... they feel like it’s one big happy family working on one application together” 3. Note on a common mix-up: Europe’s JUPITER supercomputer, sometimes assumed to be a GB200 system, in fact uses GH200 Grace Hopper Superchips, per NVIDIA’s own investor release — not GB200 4.
3.What’s inside it
SimpleStart here
Two design ideas at once: a CPU and GPU sharing memory closely enough to act like one computer (the Superchip), and up to 72 of those GPUs sharing one ultra-fast network so a giant AI model can be spread across all of them without a networking bottleneck (NVL72).
Deep diveThe technical detail
From NVIDIA’s own materials 23:
- Grace-Blackwell NVLink-C2C — the Grace CPU and its two Blackwell GPUs are connected by a 900 GB/s chip-to-chip link, closer and faster than a CPU talking to a GPU over PCIe.
- NVL72 as “one giant GPU” — Huang’s own description of the 72-GPU rack: a 130 TB/s all-to-all NVLink domain spanning the entire rack, so the 72 GPUs behave like “one big happy family working on one application together” rather than 72 separate accelerators.
- Claimed inference gains — NVIDIA states “up to 30x faster real-time trillion-parameter LLM inference” versus prior-generation (Hopper) solutions, and “up to 25x reduction in cost and energy consumption” for LLM inference workloads 12.
4.Full spec table
SimpleStart here
One Superchip: 1 Grace CPU (72 Arm cores) + 2 Blackwell GPUs, about 852GB of combined fast memory. A full NVL72 rack: 36 Superchips, 72 GPUs, 30.4 terabytes of combined fast memory 2.
Deep diveThe technical detail
| Spec | GB200 Superchip (1 Grace + 2 Blackwell) | GB200 NVL72 (36 Grace + 72 Blackwell) |
|---|
| Memory | 372GB HBM3e + up to 480GB LPDDR5X (≈852GB) | 13.4TB HBM3e + 17TB LPDDR5X (30.4TB total, NVIDIA’s own “30TB fast memory” figure) |
|---|
| GPU-to-GPU interconnect | 900 GB/s NVLink-C2C (Grace↔Blackwell) | 130 TB/s aggregate NVLink bandwidth across the 72-GPU domain |
|---|
| GPU memory bandwidth | 16 TB/s | 576 TB/s |
|---|
| NVFP4 Tensor (sparse) | 40 PFLOPS | 1,440 PFLOPS |
|---|
| FP8 Tensor | 20 PFLOPS | 720 PFLOPS |
|---|
| FP16/BF16 Tensor | 10 PFLOPS | 360 PFLOPS |
|---|
| Grace CPU cores | 72 Arm Neoverse V2 cores | 2,592 Arm Neoverse V2 cores (36×72) |
|---|
| System AI performance | — | 1.4 exaflops (NVIDIA’s own launch claim) |
|---|
| Rack power draw | Not confirmed against an NVIDIA-published figure; third-party OEM spec sheets cite roughly 120–140 kW but this is not an NVIDIA number |
|---|
| Networking | NVIDIA Quantum-X800 InfiniBand and Spectrum-X800, 800 Gb/s |
|---|
Figures per NVIDIA’s own GB200 NVL72 product page and Blackwell launch release 21. Rack power draw is explicitly marked unconfirmed because it was not found in an NVIDIA-published document, only third-party OEM specification sheets.
5.Where it’s made
SimpleStart here
The Blackwell GPU dies inside GB200 are made by TSMC on the custom 4NP process, the same as standalone B200 1. NVIDIA does not publish separate foundry detail for the Grace CPU beyond confirming the overall Blackwell platform uses TSMC.
Deep diveThe technical detail
NVIDIA’s Blackwell platform materials confirm TSMC and the 4NP process for the Blackwell GPU dies used across the platform, including inside GB200 1. No individual fab location is published for either the Grace CPU or the Blackwell GPU dies.
6.Which systems use it
SimpleStart here
NVIDIA’s own reference rack is the DGX GB200 NVL72 / DGX SuperPOD. Oracle Cloud Infrastructure and Microsoft Azure have both announced GB200/GB300-class deployments on their own blogs 56.
Deep diveThe technical detail
NVIDIA’s own DGX SuperPOD reference architecture is built around GB200 NVL72 racks. Oracle Cloud Infrastructure’s own blog describes scaling “NVIDIA GB200 NVL72 deployments” with dedicated OCI APIs 5. Microsoft Azure’s own NVIDIA-blog post describes what it calls the “world’s first” GB300 NVL72 (the Blackwell Ultra successor generation) supercomputing cluster for OpenAI, at 4,608 GPUs 6 — direct lineage from GB200 but a later generation, included here because it is the clearest publicly documented hyperscale deployment of this rack design. As noted above, JUPITER (Europe’s planned first exascale system) uses GH200 Grace Hopper Superchips, not GB200 4.
7.Official pricing
SimpleStart here
NVIDIA has not published a unit price for a GB200 Superchip or a full NVL72 rack. CoreWeave, an official cloud partner, lists “NVIDIA GB200 NVL72 (4 GPUs)” — priced per its NVL4 compute-tray building block — at $42.00/hr on-demand, or $10.50/hr for a single-GPU inference rate 7.
Deep diveThe technical detail
CoreWeave’s “4 GPUs” unit corresponds to NVIDIA’s own GB200 NVL4 building-block configuration (2 Grace CPUs + 4 Blackwell GPUs), a real, separately documented NVIDIA configuration, not a pricing error 7. Oracle Cloud Infrastructure and Microsoft Azure have both announced GB200/GB300 NVL72 availability on their own blogs 56, but neither page’s exact hourly or unit price is reproduced here. No NVIDIA-published price for a full rack exists in the materials reviewed.
9.What came before, what came next
SimpleStart here
GB200 has no direct 1:1 predecessor — it is NVIDIA’s first Grace-plus-dual-Blackwell Superchip. The closest conceptual predecessor is the GH200 Grace Hopper Superchip (Grace + a single H200 GPU). Successor: GB300 “Grace Blackwell Ultra”, announced 18 March 2025 8.
Deep diveThe technical detail
NVIDIA’s own Blackwell Ultra announcement states the GB300 NVL72 configuration delivers “1.5x more AI performance than GB200 NVL72” 8, making the lineage explicit: GH200 (Grace Hopper, conceptual predecessor) → GB200 (Grace Blackwell) → GB300 (Grace Blackwell Ultra).
10.Hidden in plain sight
SimpleStart here
NVIDIA’s own CEO describes the entire 72-GPU, liquid-cooled rack as “one giant GPU,” not 72 separate accelerators — a direct quote from the GTC 2024 keynote 3, reflecting the 130 TB/s NVLink domain spanning the whole rack.
Deep diveThe technical detail
The framing matters technically, not just rhetorically: software running across the NVL72 rack addresses the 72 GPUs and their 30.4TB of combined memory through the same NVLink domain, rather than treating each server as a separate unit connected by conventional networking. This is what makes NVIDIA’s “one giant GPU” description 3 a claim about memory architecture and interconnect topology, not marketing shorthand for “a lot of GPUs in one place.”