How to read this page. Each section starts Simple, then goes to a Deep dive: stop wherever you have what you need. The small numbers are sources: click one to open the original document. Where the maker has not published a figure -- a price, a die size, a factory address -- this page says so rather than estimate it.
1.At a glance
SimpleStart here
WSE-3, announced 13 March 2024, is Cerebras’s current wafer-sized chip: 4 trillion transistors and 900,000 cores on one piece of silicon, delivering 125 petaflops of AI compute 1. It is the chip behind Cerebras Inference, launched August 2024, which the company calls the world’s fastest AI inference cloud 2.
Deep diveThe technical detail
Cerebras’s own head-to-head math against NVIDIA’s then-new B200: “a single CS-3 equates to about 3.5 DGX B200 servers” on raw compute (125 vs. 36 petaflops), a 2.2x improvement in performance per watt 3 — see the performance section below for why that is not the same as CS-3 drawing less power in absolute terms. This is Cerebras’s own comparison, not an independent benchmark.
2.Launch and history
SimpleStart here
WSE-3 is the generation where Cerebras’s business model changed. It launched the same day as a new supercomputer deal (Condor Galaxy 3, with G42) but within five months had also launched a pay-per-token inference cloud — the product line that now earns Cerebras more revenue than selling systems outright (see the Cerebras brand guide on this site) 1 4 2.
Deep diveThe technical detail
At its inference launch, Cerebras served Meta’s Llama 3.1 8B model at “1,800 tokens a second” and Llama 3.1 70B at “450 tokens a second” — speeds it attributed directly to keeping the model’s weights on the wafer’s own memory rather than fetching them from separate memory chips 2 5.
3.What’s inside it
SimpleStart here
900,000 cores, up from 850,000 on WSE-2, joined by an upgraded “SwarmX” interconnect that links up to 2,048 CS-3 systems into one cluster 1 — a “10x improvement” over WSE-2’s 192-system limit, per Cerebras’s own blog 6.
Deep diveThe technical detail
For models too large for one wafer’s on-chip memory, Cerebras pairs the chip with an external unit called MemoryX, offered in sizes from 24 TB and 36 TB (enterprise) up to 120 TB and 1,200 TB (hyperscale); the largest configuration can hold “models with 24 trillion parameters” 6. Cerebras says a full 2,048-system cluster, using its Weight Streaming architecture, “looks and programs like a single chip” 6.
4.Full spec table
SimpleStart here
| Spec | Cerebras WSE-3 (CS-3) |
|---|
| Transistors | 4 trillion 1 |
|---|
| Cores | 900,000 AI cores 1 |
|---|
| On-chip memory | 44 GB SRAM 1 |
|---|
| Memory bandwidth | 21 petabytes/s, Cerebras says — “7,000x that of an H100” 5 |
|---|
| AI compute | 125 petaflops 1 |
|---|
| External memory (MemoryX) | Up to 1.2 PB (see note) 1 |
|---|
| Process node | TSMC 5 nm 1 |
|---|
| System size / power | 15U, up to 23 kW 3 |
|---|
Cerebras’s own launch materials give two different sets of MemoryX tier sizes: the March 2024 announcement lists “1.5TB, 12TB, or 1.2PB” options 1, while a separate Cerebras blog post lists “24TB and 36TB” for enterprise customers and “120TB and 1,200TB” for hyperscalers 6. Only the top tier (1.2 PB = 1,200 TB) matches between the two; this page reports both rather than picking one, since Cerebras’s own materials do not reconcile the difference.
Deep diveThe technical detail
Cerebras’s own comparison against a full NVIDIA NVL72 rack (72 GPUs): a single CS-3 “provides more than 200x the interconnect bandwidth” — stated as 27 petabytes per second of aggregate on-wafer fabric bandwidth across the chip’s 900,000 cores — and “over 80x the memory capacity” when paired with its largest MemoryX configuration 3. This 27 PB/s core-to-core fabric figure is a different measurement from the 21 PB/s figure above, which Cerebras states separately as memory-to-core bandwidth — both describe traffic within one wafer, not between separate chips, and Cerebras’s own materials do not explain how the two relate to each other.
5.Where it’s made
SimpleStart here
Designed by Cerebras; the wafer made by TSMC on its 5 nm process, its most advanced node yet for this chip line 1.
Deep diveThe technical detail
Cerebras’s January 2025 blog post on this generation’s manufacturing says each wafer builds 970,000 physical cores and activates 900,000 of them, routing around the rest, and describes the chip as “about 100x more fault tolerant” than a GPU because a single core failure costs a much smaller share of the total wafer 7.
6.Which systems use it
SimpleStart here
Sold as the CS-3 system and rented through Cerebras’s own cloud. Named deployments include Condor Galaxy 3 with G42 (announced the same day as the chip itself) and a March 2026 collaboration to run CS-3 systems inside AWS’s own data centres 4 8.
Deep diveThe technical detail
By March 2025 Cerebras said its own data centres in Santa Clara, Stockton and Dallas were already online, running on this generation, with more sites in the US and Europe planned 9. Under the AWS collaboration, Cerebras’s CS-3 handles one stage of an AI request and AWS’s own Trainium chips handle another (see the Amazon Trainium guide on this site) 8.
7.Official pricing
SimpleStart here
Cerebras has not published a CS-3 hardware list price. Its cloud service, built on CS-3, bills per token — see the Cerebras brand guide on this site for current rates, which apply to the models served on this chip.
Deep diveThe technical detail
At its August 2024 inference launch, Cerebras priced Llama 3.1 8B at 10 cents and Llama 3.1 70B at 60 cents per million tokens — the earliest published Cerebras Inference prices, all served from CS-3 hardware 2.
9.What came before, what came next
SimpleStart here
Came before: WSE-2 (CS-2), covered on its own page on this site.
Came after: WSE-3 Turbo and CS-4, announced August 2026 but not yet confirmed shipped as of this page’s last check — see the Cerebras brand guide on this site.
Deep diveThe technical detail
WSE-3 launched almost three years after WSE-2 (April 2021 to March 2024), a longer gap than the roughly 20 months between the first two generations 10 1. Its successor, CS-4, was announced in August 2026 with Cerebras saying “first CS-4 shipments begin this quarter” — language this site treats as a projection, not a confirmed shipment, until an actual customer delivery is reported.
10.Hidden in plain sight
SimpleStart here
Cerebras’s own numbers show this chip is not actually faster per core than its predecessor — it wins by adding cores and memory, not by making each core quicker: cores grew from 850,000 to 900,000 (about 6%), while the compute figure Cerebras publishes (125 petaflops) is a whole-wafer number, not a per-core speed 10 1.
Deep diveThe technical detail
The August 2024 inference launch reframed this chip’s purpose in Cerebras’s own business: WSE-3 was announced in March 2024 as a training chip, competing on raw petaflops against NVIDIA’s training GPUs 1 3, but within five months Cerebras was marketing the same silicon’s speed advantage for serving finished models to users, not just training them 2 — the same hardware, sold against two different competitive claims within one calendar year.