AI Chips · Product page

Cerebras CS-3 (WSE-3): The Complete Guide

The chip Cerebras still sells and rents today — 4 trillion transistors on one wafer, clustered up to 2,048 at a time, and the machine behind its instant-inference cloud. Simple to expert, every number links to its source.

Stuck at any point? Ask an AI about this page →
Ask an AI about this chip:CAGPX
Cerebras CS-3 (WSE-3) 10 sections · 2 levels 10 linked sources Checked September 2026
How to read this page. Each section starts Simple, then goes to a Deep dive: stop wherever you have what you need. The small numbers are sources: click one to open the original document. Where the maker has not published a figure -- a price, a die size, a factory address -- this page says so rather than estimate it.

1.At a glance

SimpleStart here

WSE-3, announced 13 March 2024, is Cerebras’s current wafer-sized chip: 4 trillion transistors and 900,000 cores on one piece of silicon, delivering 125 petaflops of AI compute 1. It is the chip behind Cerebras Inference, launched August 2024, which the company calls the world’s fastest AI inference cloud 2.

Deep diveThe technical detail

Cerebras’s own head-to-head math against NVIDIA’s then-new B200: “a single CS-3 equates to about 3.5 DGX B200 servers” on raw compute (125 vs. 36 petaflops), a 2.2x improvement in performance per watt 3 — see the performance section below for why that is not the same as CS-3 drawing less power in absolute terms. This is Cerebras’s own comparison, not an independent benchmark.

2.Launch and history

SimpleStart here

WSE-3 is the generation where Cerebras’s business model changed. It launched the same day as a new supercomputer deal (Condor Galaxy 3, with G42) but within five months had also launched a pay-per-token inference cloud — the product line that now earns Cerebras more revenue than selling systems outright (see the Cerebras brand guide on this site) 1 4 2.

Deep diveThe technical detail

At its inference launch, Cerebras served Meta’s Llama 3.1 8B model at “1,800 tokens a second” and Llama 3.1 70B at “450 tokens a second” — speeds it attributed directly to keeping the model’s weights on the wafer’s own memory rather than fetching them from separate memory chips 2 5.

3.What’s inside it

SimpleStart here

900,000 cores, up from 850,000 on WSE-2, joined by an upgraded “SwarmX” interconnect that links up to 2,048 CS-3 systems into one cluster 1 — a “10x improvement” over WSE-2’s 192-system limit, per Cerebras’s own blog 6.

Deep diveThe technical detail

For models too large for one wafer’s on-chip memory, Cerebras pairs the chip with an external unit called MemoryX, offered in sizes from 24 TB and 36 TB (enterprise) up to 120 TB and 1,200 TB (hyperscale); the largest configuration can hold “models with 24 trillion parameters” 6. Cerebras says a full 2,048-system cluster, using its Weight Streaming architecture, “looks and programs like a single chip” 6.

4.Full spec table

SimpleStart here
SpecCerebras WSE-3 (CS-3)
Transistors4 trillion 1
Cores900,000 AI cores 1
On-chip memory44 GB SRAM 1
Memory bandwidth21 petabytes/s, Cerebras says — “7,000x that of an H100” 5
AI compute125 petaflops 1
External memory (MemoryX)Up to 1.2 PB (see note) 1
Process nodeTSMC 5 nm 1
System size / power15U, up to 23 kW 3

Cerebras’s own launch materials give two different sets of MemoryX tier sizes: the March 2024 announcement lists “1.5TB, 12TB, or 1.2PB” options 1, while a separate Cerebras blog post lists “24TB and 36TB” for enterprise customers and “120TB and 1,200TB” for hyperscalers 6. Only the top tier (1.2 PB = 1,200 TB) matches between the two; this page reports both rather than picking one, since Cerebras’s own materials do not reconcile the difference.

Deep diveThe technical detail

Cerebras’s own comparison against a full NVIDIA NVL72 rack (72 GPUs): a single CS-3 “provides more than 200x the interconnect bandwidth” — stated as 27 petabytes per second of aggregate on-wafer fabric bandwidth across the chip’s 900,000 cores — and “over 80x the memory capacity” when paired with its largest MemoryX configuration 3. This 27 PB/s core-to-core fabric figure is a different measurement from the 21 PB/s figure above, which Cerebras states separately as memory-to-core bandwidth — both describe traffic within one wafer, not between separate chips, and Cerebras’s own materials do not explain how the two relate to each other.

5.Where it’s made

SimpleStart here

Designed by Cerebras; the wafer made by TSMC on its 5 nm process, its most advanced node yet for this chip line 1.

Deep diveThe technical detail

Cerebras’s January 2025 blog post on this generation’s manufacturing says each wafer builds 970,000 physical cores and activates 900,000 of them, routing around the rest, and describes the chip as “about 100x more fault tolerant” than a GPU because a single core failure costs a much smaller share of the total wafer 7.

6.Which systems use it

SimpleStart here

Sold as the CS-3 system and rented through Cerebras’s own cloud. Named deployments include Condor Galaxy 3 with G42 (announced the same day as the chip itself) and a March 2026 collaboration to run CS-3 systems inside AWS’s own data centres 4 8.

Deep diveThe technical detail

By March 2025 Cerebras said its own data centres in Santa Clara, Stockton and Dallas were already online, running on this generation, with more sites in the US and Europe planned 9. Under the AWS collaboration, Cerebras’s CS-3 handles one stage of an AI request and AWS’s own Trainium chips handle another (see the Amazon Trainium guide on this site) 8.

7.Official pricing

SimpleStart here

Cerebras has not published a CS-3 hardware list price. Its cloud service, built on CS-3, bills per token — see the Cerebras brand guide on this site for current rates, which apply to the models served on this chip.

Deep diveThe technical detail

At its August 2024 inference launch, Cerebras priced Llama 3.1 8B at 10 cents and Llama 3.1 70B at 60 cents per million tokens — the earliest published Cerebras Inference prices, all served from CS-3 hardware 2.

8.Real-world performance

SimpleStart here

Cerebras’s own comparison against NVIDIA’s DGX B200: 125 petaflops versus 36 petaflops — about 3.5 DGX B200 servers’ worth of compute in one CS-3 — which Cerebras calls “a 2.2x improvement in performance per watt” 3.

Deep diveThe technical detail

This is Cerebras’s own published comparison, not an independent third-party benchmark, and its “half the power consumption” framing needs a caveat. By the same source’s own figures, CS-3 draws up to 23 kW peak against a DGX B200’s 14.3 kW — more absolute power, not less. “Half” refers to power per unit of compute (performance per watt), which is where the 2.2x figure comes from; it is not a claim that a CS-3 system pulls fewer watts off the wall 3. Cerebras also claims a full 2,048-system CS-3 cluster can “train Llama2-70B from scratch in less than a day” versus “approximately a month” it attributes to Meta’s own GPU infrastructure for the same task 6 — a claim about a specific competitor’s reported training time, not a controlled side-by-side test Cerebras ran itself.

9.What came before, what came next

SimpleStart here
Came before: WSE-2 (CS-2), covered on its own page on this site. Came after: WSE-3 Turbo and CS-4, announced August 2026 but not yet confirmed shipped as of this page’s last check — see the Cerebras brand guide on this site.
Deep diveThe technical detail

WSE-3 launched almost three years after WSE-2 (April 2021 to March 2024), a longer gap than the roughly 20 months between the first two generations 10 1. Its successor, CS-4, was announced in August 2026 with Cerebras saying “first CS-4 shipments begin this quarter” — language this site treats as a projection, not a confirmed shipment, until an actual customer delivery is reported.

10.Hidden in plain sight

SimpleStart here

Cerebras’s own numbers show this chip is not actually faster per core than its predecessor — it wins by adding cores and memory, not by making each core quicker: cores grew from 850,000 to 900,000 (about 6%), while the compute figure Cerebras publishes (125 petaflops) is a whole-wafer number, not a per-core speed 10 1.

Deep diveThe technical detail

The August 2024 inference launch reframed this chip’s purpose in Cerebras’s own business: WSE-3 was announced in March 2024 as a training chip, competing on raw petaflops against NVIDIA’s training GPUs 1 3, but within five months Cerebras was marketing the same silicon’s speed advantage for serving finished models to users, not just training them 2 — the same hardware, sold against two different competitive claims within one calendar year.

11.Sources

10 sources, checked September 2026. Where NVIDIA or another maker has not published a figure, this page says so rather than estimate it.

  1. Cerebras: Cerebras announces third-generation Wafer Scale Engine (13 Mar 2024)Official
  2. Cerebras: Cerebras launches the world’s fastest AI inference (27 Aug 2024)Official
  3. Cerebras Blog: Cerebras CS-3 vs NVIDIA B200: 2024 AI accelerators compared (Apr 2024)Official
  4. Cerebras: Cerebras and G42 break ground on Condor Galaxy 3, an 8 exaFLOPs AI supercomputer (13 Mar 2024)Official
  5. Cerebras Blog: Introducing Cerebras Inference: AI at instant speed (Aug 2024)Official
  6. Cerebras Blog: Cerebras CS-3: the world’s fastest and most scalable AI acceleratorOfficial
  7. Cerebras Blog: 100x defect tolerance: how Cerebras solved the yield problem (13 Jan 2025)Official
  8. Cerebras: AWS collaboration: CS-3 in AWS data centres (13 Mar 2026)Official
  9. Cerebras: Cerebras announces six new AI data centres across North America and Europe (11 Mar 2025)Official
  10. Cerebras: Second-generation Wafer Scale Engine, 2.6 trillion transistors (20 Apr 2021)Official

Ask an AI about this page

Opens your assistant with this page as the source, and a question rather than a summary. It will ask what you are building before it answers.

ChatGPTClaudeGeminiPerplexityGrok

Nothing is sent from here. The link carries only this page’s title and address.