How to read this page. Each section starts Simple, then goes to a Deep dive: stop wherever you have what you need. The small numbers are sources: click one to open the original document. Where the maker has not published a figure -- a price, a die size, a factory address -- this page says so rather than estimate it.
1.At a glance
SimpleStart here
The B200 is NVIDIA’s Blackwell-architecture data-centre GPU, announced 18 March 2024 at GTC 1. It is the first NVIDIA data-centre GPU built from two separate dies fused together that the software sees as a single chip.
Deep diveThe technical detail
Architecture Blackwell (second generation). NVIDIA’s Q3 fiscal 2025 earnings release (filed with the SEC 20 November 2024) has Jensen Huang stating Blackwell is “in full production” 2, with shipments ramping into late 2024 and early 2025. Ships as the SXM form factor in HGX B200 8-GPU baseboards 3.
2.Launch and history
SimpleStart here
NVIDIA announced Blackwell as a full computing platform, not just a chip, at GTC 2024. Jensen Huang described the largest rack-scale configuration built around it as functioning like “one giant GPU” rather than a collection of separate chips.
Deep diveThe technical detail
Announced 18 March 2024 as part of the “Blackwell Platform” launch 1, with availability stated as “starting later this year” (2024) 1. NVIDIA’s own Q3 FY2025 SEC filing quotes Huang confirming full production status by the November 2024 earnings call 2. At the GTC 2024 keynote, describing the largest Blackwell rack-scale system, Huang said: “one giant GPU” — “a multi-node, liquid-cooled, rack-scale system” 4.
3.What’s inside it
SimpleStart here
A single B200 “chip” is actually two reticle-limit dies connected by an extremely fast internal link, working together as one GPU the software never has to treat as two separate devices 1.
Deep diveThe technical detail
From NVIDIA’s own materials 15:
- Dual-die design — NVIDIA describes B200 as “two-reticle limit GPU dies connected by a 10 TB/second chip-to-chip link,” presented to software as a single unified GPU.
- Second-generation Transformer Engine — adds FP4 support via “micro-tensor scaling,” giving “double the compute and model sizes compared to predecessor.”
- Confidential computing — NVIDIA’s Blackwell architecture page describes hardware-based confidential computing with “nearly identical throughput performance compared to unencrypted modes.”
- Dedicated RAS Engine — provides “AI-powered predictive-management capabilities” that “continuously monitor thousands of data points” for reliability.
4.Full spec table
SimpleStart here
208 billion transistors across the two fused dies, 180GB of HBM3e memory, and a chip-to-chip link between the two halves running at 10 TB/s 6.
Deep diveThe technical detail
| Spec | B200 SXM |
|---|
| Process node | TSMC 4NP (custom) 5 |
|---|
| Foundry | TSMC 5 |
|---|
| Transistors | 208 billion 1 |
|---|
| Die configuration | Two-reticle limit GPU dies, 10 TB/second chip-to-chip link (NV-HBI), one CUDA-visible GPU 1 |
|---|
| Die size | Not published by NVIDIA |
|---|
| CUDA / Tensor core count | Not published by NVIDIA for Blackwell data-centre parts 6 |
|---|
| Memory | 180GB HBM3e |
|---|
| Memory bandwidth | 7.7 TB/s |
|---|
| TDP | Configurable up to 1,000 W (SXM) |
|---|
| FP64 / FP32 | 37 TFLOPS / 75 TFLOPS |
|---|
| TF32 Tensor (sparse) | 2.2 PFLOPS |
|---|
| FP16/BF16 Tensor (sparse) | 4.5 PFLOPS |
|---|
| FP8/FP6 Tensor (sparse) | 9 PFLOPS |
|---|
| INT8 Tensor (sparse) | 9 POPS |
|---|
| FP4 Tensor (sparse) | 18 PFLOPS (dense = half, per NVIDIA’s own footnote) |
|---|
| NVLink | 5th gen, 1.8 TB/s bidirectional per GPU; scales to 576 GPUs |
|---|
| Form factor | SXM (HGX B200 8-GPU baseboard: 1,440GB total memory, 64 TB/s aggregate HBM3e bandwidth, ~14.3kW system TDP) |
|---|
Figures per NVIDIA’s own Blackwell datasheet and architecture page 65. Die size and CUDA/Tensor core counts are marked “not published” because NVIDIA has stopped disclosing them for Blackwell data-centre parts, unlike the H100 whitepaper which did.
5.Where it’s made
SimpleStart here
B200 is made by TSMC, on a custom process NVIDIA calls 4NP 5. NVIDIA designs the chip; TSMC manufactures it.
Deep diveThe technical detail
NVIDIA’s own Blackwell architecture materials name TSMC and the 4NP process for both dies that make up a B200 56. No exact die size or individual fab location is published by either company.
6.Which systems use it
SimpleStart here
NVIDIA’s own system is DGX B200 (8 chips) 3. HGX B200 is the OEM baseboard adopted by Dell, HPE, Lenovo, Cisco, Supermicro, ASUS and GIGABYTE, per NVIDIA’s own partner list 1.
Deep diveThe technical detail
DGX B200 is NVIDIA’s reference 8-GPU system 3. HGX B200 is the rack-scale building block OEMs build servers around, with NVIDIA naming Dell, HPE, Lenovo, Cisco, Supermicro, ASUS and GIGABYTE as adopting partners 1. NVIDIA also documents a DGX SuperPOD reference architecture built on B200.
7.Official pricing
SimpleStart here
NVIDIA has not published a unit or system price for B200. CoreWeave, an official cloud partner, lists HGX B200 (8 GPUs) at $68.80/hr on-demand, $34.11/hr spot, or $8.60/hr for a single-GPU inference instance 7.
Deep diveThe technical detail
CoreWeave’s own pricing page is the only source found in this page’s research pass that publishes a specific, checkable B200 rate 7. No NVIDIA-published unit or system price for B200 or HGX B200 exists in the materials reviewed — this is marked as not published rather than estimated.
9.What came before, what came next
SimpleStart here
Came before: H100 / H200 (Hopper). Came after: B300, marketed as “Blackwell Ultra”, announced 18 March 2025 8.
Deep diveThe technical detail
B200 succeeds the Hopper generation (H100/H200). NVIDIA announced its direct successor, B300 (“Blackwell Ultra”), on 18 March 2025, claiming “1.5x more AI performance than GB200 NVL72” in the GB300 NVL72 configuration, with availability stated as “second half of 2025” 8. The next full platform generation after Blackwell is Rubin, which NVIDIA has announced but which is not independently re-confirmed via a primary press release on this page.
10.Hidden in plain sight
SimpleStart here
A “single” B200 chip is not one piece of silicon. NVIDIA’s own description is “two-reticle limit GPU dies” fused by an internal 10 TB/second link that the software sees as one GPU — the first time NVIDIA has shipped a data-centre GPU built this way 1.
Deep diveThe technical detail
Every prior NVIDIA data-centre GPU on this site (V100, A100, H100, H200) is a single monolithic die. B200 is the first exception: NVIDIA states it is “two-reticle limit GPU dies connected by a 10 TB/second chip-to-chip link” 1, and CUDA code addresses the pair as one GPU with no awareness of the split. NVIDIA has also stopped publishing a die size or CUDA-core count for this generation, a level of disclosure every prior generation on this site provided.