AI Chips · Product page

NVIDIA GB200 (Grace Blackwell Superchip): The Complete Guide

One CPU and two GPUs on a single board, scaled up to a 72-GPU rack NVIDIA calls “one giant GPU.” Simple to expert, every number links to its source.

Ask an AI about this chip:CAGPX
NVIDIA GB200 10 sections · 2 levels 8 linked sources Checked September 2026
How to read this page. Each section starts Simple, then goes to a Deep dive: stop wherever you have what you need. The small numbers are sources: click one to open the original document. Where the maker has not published a figure -- a price, a die size, a factory address -- this page says so rather than estimate it.

1.At a glance

SimpleStart here

The GB200 is not one chip — it is a “Superchip”: one NVIDIA Grace CPU plus two Blackwell GPUs on a single board, announced 18 March 2024 alongside B200 1. NVIDIA sells it both as an individual Superchip and scaled into a 72-GPU rack called the GB200 NVL72.

Deep diveThe technical detail

Design: 1× Grace (Arm CPU) + 2× Blackwell GPU, connected by NVLink-C2C. The NVL72 rack combines 36 Grace CPUs and 72 Blackwell GPUs into one NVLink domain 2. This is NVIDIA’s first rack-scale product marketed as behaving like a single accelerator rather than a cluster of servers.

2.Launch and history

SimpleStart here

GB200 exists because a rack full of separate GPU servers, however fast individually, hits a wall when a model is too large to fit in one server’s memory and has to be split across many machines with slower networking between them. NVIDIA’s answer was to fuse CPU, GPU and rack-scale networking into one design from the start.

Deep diveThe technical detail

Announced at GTC on 18 March 2024, alongside B200 12. At the same keynote, describing the NVL72 rack, Jensen Huang said: “one giant GPU” — “a multi-node, liquid-cooled, rack-scale system... they feel like it’s one big happy family working on one application together” 3. Note on a common mix-up: Europe’s JUPITER supercomputer, sometimes assumed to be a GB200 system, in fact uses GH200 Grace Hopper Superchips, per NVIDIA’s own investor release — not GB200 4.

3.What’s inside it

SimpleStart here

Two design ideas at once: a CPU and GPU sharing memory closely enough to act like one computer (the Superchip), and up to 72 of those GPUs sharing one ultra-fast network so a giant AI model can be spread across all of them without a networking bottleneck (NVL72).

Deep diveThe technical detail

From NVIDIA’s own materials 23:

  • Grace-Blackwell NVLink-C2C — the Grace CPU and its two Blackwell GPUs are connected by a 900 GB/s chip-to-chip link, closer and faster than a CPU talking to a GPU over PCIe.
  • NVL72 as “one giant GPU” — Huang’s own description of the 72-GPU rack: a 130 TB/s all-to-all NVLink domain spanning the entire rack, so the 72 GPUs behave like “one big happy family working on one application together” rather than 72 separate accelerators.
  • Claimed inference gains — NVIDIA states “up to 30x faster real-time trillion-parameter LLM inference” versus prior-generation (Hopper) solutions, and “up to 25x reduction in cost and energy consumption” for LLM inference workloads 12.

4.Full spec table

SimpleStart here

One Superchip: 1 Grace CPU (72 Arm cores) + 2 Blackwell GPUs, about 852GB of combined fast memory. A full NVL72 rack: 36 Superchips, 72 GPUs, 30.4 terabytes of combined fast memory 2.

Deep diveThe technical detail
SpecGB200 Superchip (1 Grace + 2 Blackwell)GB200 NVL72 (36 Grace + 72 Blackwell)
Memory372GB HBM3e + up to 480GB LPDDR5X (≈852GB)13.4TB HBM3e + 17TB LPDDR5X (30.4TB total, NVIDIA’s own “30TB fast memory” figure)
GPU-to-GPU interconnect900 GB/s NVLink-C2C (Grace↔Blackwell)130 TB/s aggregate NVLink bandwidth across the 72-GPU domain
GPU memory bandwidth16 TB/s576 TB/s
NVFP4 Tensor (sparse)40 PFLOPS1,440 PFLOPS
FP8 Tensor20 PFLOPS720 PFLOPS
FP16/BF16 Tensor10 PFLOPS360 PFLOPS
Grace CPU cores72 Arm Neoverse V2 cores2,592 Arm Neoverse V2 cores (36×72)
System AI performance—1.4 exaflops (NVIDIA’s own launch claim)
Rack power drawNot confirmed against an NVIDIA-published figure; third-party OEM spec sheets cite roughly 120–140 kW but this is not an NVIDIA number
NetworkingNVIDIA Quantum-X800 InfiniBand and Spectrum-X800, 800 Gb/s

Figures per NVIDIA’s own GB200 NVL72 product page and Blackwell launch release 21. Rack power draw is explicitly marked unconfirmed because it was not found in an NVIDIA-published document, only third-party OEM specification sheets.

5.Where it’s made

SimpleStart here

The Blackwell GPU dies inside GB200 are made by TSMC on the custom 4NP process, the same as standalone B200 1. NVIDIA does not publish separate foundry detail for the Grace CPU beyond confirming the overall Blackwell platform uses TSMC.

Deep diveThe technical detail

NVIDIA’s Blackwell platform materials confirm TSMC and the 4NP process for the Blackwell GPU dies used across the platform, including inside GB200 1. No individual fab location is published for either the Grace CPU or the Blackwell GPU dies.

6.Which systems use it

SimpleStart here

NVIDIA’s own reference rack is the DGX GB200 NVL72 / DGX SuperPOD. Oracle Cloud Infrastructure and Microsoft Azure have both announced GB200/GB300-class deployments on their own blogs 56.

Deep diveThe technical detail

NVIDIA’s own DGX SuperPOD reference architecture is built around GB200 NVL72 racks. Oracle Cloud Infrastructure’s own blog describes scaling “NVIDIA GB200 NVL72 deployments” with dedicated OCI APIs 5. Microsoft Azure’s own NVIDIA-blog post describes what it calls the “world’s first” GB300 NVL72 (the Blackwell Ultra successor generation) supercomputing cluster for OpenAI, at 4,608 GPUs 6 — direct lineage from GB200 but a later generation, included here because it is the clearest publicly documented hyperscale deployment of this rack design. As noted above, JUPITER (Europe’s planned first exascale system) uses GH200 Grace Hopper Superchips, not GB200 4.

7.Official pricing

SimpleStart here

NVIDIA has not published a unit price for a GB200 Superchip or a full NVL72 rack. CoreWeave, an official cloud partner, lists “NVIDIA GB200 NVL72 (4 GPUs)” — priced per its NVL4 compute-tray building block — at $42.00/hr on-demand, or $10.50/hr for a single-GPU inference rate 7.

Deep diveThe technical detail

CoreWeave’s “4 GPUs” unit corresponds to NVIDIA’s own GB200 NVL4 building-block configuration (2 Grace CPUs + 4 Blackwell GPUs), a real, separately documented NVIDIA configuration, not a pricing error 7. Oracle Cloud Infrastructure and Microsoft Azure have both announced GB200/GB300 NVL72 availability on their own blogs 56, but neither page’s exact hourly or unit price is reproduced here. No NVIDIA-published price for a full rack exists in the materials reviewed.

8.Real-world performance

SimpleStart here

NVIDIA’s own launch claims for GB200 NVL72: up to 30x faster real-time trillion-parameter LLM inference and up to 25x lower cost and energy use versus prior-generation Hopper solutions 12 — company-stated, system-level figures.

Deep diveThe technical detail

These multipliers are NVIDIA’s own comparisons of an entire NVL72 rack against an unspecified prior-generation Hopper-based configuration, not a benchmark of a single chip in isolation, and not an independently reproduced third-party result. The rated system AI performance for a full rack is stated by NVIDIA as 1.4 exaflops 1.

9.What came before, what came next

SimpleStart here

GB200 has no direct 1:1 predecessor — it is NVIDIA’s first Grace-plus-dual-Blackwell Superchip. The closest conceptual predecessor is the GH200 Grace Hopper Superchip (Grace + a single H200 GPU). Successor: GB300 “Grace Blackwell Ultra”, announced 18 March 2025 8.

Deep diveThe technical detail

NVIDIA’s own Blackwell Ultra announcement states the GB300 NVL72 configuration delivers “1.5x more AI performance than GB200 NVL72” 8, making the lineage explicit: GH200 (Grace Hopper, conceptual predecessor) → GB200 (Grace Blackwell) → GB300 (Grace Blackwell Ultra).

10.Hidden in plain sight

SimpleStart here

NVIDIA’s own CEO describes the entire 72-GPU, liquid-cooled rack as “one giant GPU,” not 72 separate accelerators — a direct quote from the GTC 2024 keynote 3, reflecting the 130 TB/s NVLink domain spanning the whole rack.

Deep diveThe technical detail

The framing matters technically, not just rhetorically: software running across the NVL72 rack addresses the 72 GPUs and their 30.4TB of combined memory through the same NVLink domain, rather than treating each server as a separate unit connected by conventional networking. This is what makes NVIDIA’s “one giant GPU” description 3 a claim about memory architecture and interconnect topology, not marketing shorthand for “a lot of GPUs in one place.”

11.Sources

8 sources, checked September 2026. Where NVIDIA or another maker has not published a figure, this page says so rather than estimate it.

  1. NVIDIA Newsroom: NVIDIA Blackwell Platform Arrives to Power a New Era of ComputingOfficial
  2. NVIDIA: GB200 NVL72 product pageOfficial
  3. NVIDIA Blog: “We Created a Processor for the Generative AI Era,” NVIDIA CEO SaysOfficial
  4. NVIDIA Investor Relations: NVIDIA Powers Europe’s Fastest SupercomputerOfficial
  5. Oracle Cloud Infrastructure Blog: Behind the Scenes: Scale your NVIDIA GB200 NVL72 deployments with dedicated OCI APIsOfficial
  6. NVIDIA Blog: Microsoft Azure Unveils World’s First NVIDIA GB300 NVL72 Supercomputing Cluster for OpenAIOfficial
  7. CoreWeave: Cloud GPU PricingOfficial
  8. NVIDIA Newsroom: NVIDIA Blackwell Ultra AI Factory Platform Paves Way for Age of AI ReasoningOfficial