How to read this page. Each section starts Simple, then goes to a Deep dive: stop wherever you have what you need. The small numbers are sources: click one to open the original document. Where the maker has not published a figure -- a price, a die size, a factory address -- this page says so rather than estimate it.
1.At a glance
SimpleStart here
TPU v3 ran so hot that Google plumbed liquid cooling into its data centers for the first time — a first Google’s own retrospective calls out by name 1. Announced and shown entering alpha around Google I/O and Google Cloud Next ’18.
Deep diveThe technical detail
Google Cloud Next ’18 (24–26 July 2018) confirmed “Cloud TPU Pods and TPU v3 are now available in alpha” 2. The exact figure Google gave live on stage at I/O for peak pod performance is not independently confirmed in an official Google blog post; this page uses Google’s current documented spec instead of the press-reported keynote number.
2.Launch and history
SimpleStart here
TPU v3 pushed enough compute into each chip that air cooling could no longer keep up. Google’s own infrastructure team says liquid cooling let it double chip density and quadruple the size of its AI supercomputer compared to the air-cooled generation before it.
Deep diveThe technical detail
Google’s own words: “Google first used liquid cooling in TPU v3 that was deployed in 2018,” and that switch let Google “double chip density” and build a liquid-cooled supercomputer “four times” the size of its air-cooled TPU v2 predecessor 3. This is confirmed at two separate official Google Cloud sources 13.
3.What’s inside it
SimpleStart here
TPU v3 moved to HBM2 memory and a 2D torus network connecting up to 1,024 chips into one pod 4 — a step up in both memory technology and scale from TPU v2.
Deep diveThe technical detail
Per Google’s own current documentation 4: 32 GiB of HBM2 per chip at 900 GB/s, a 2D torus interconnect topology, 340 TB/s of pod-wide all-reduce bandwidth and 6.4 TB/s of bisection bandwidth. Google does not publish a process node for TPU v3.
4.Full spec table
SimpleStart here
123 teraflops per chip, 32GB of HBM2 memory, and a 1,024-chip pod delivering 126 petaflops 4.
Deep diveThe technical detail
| Spec | Google TPU v3 |
|---|
| Peak compute (per chip, bf16) | 123 teraflops |
|---|
| Peak compute (per pod, bf16) | 126 petaflops |
|---|
| Memory | 32 GiB HBM2 per chip, 900 GB/s |
|---|
| Pod size | 1,024 chips, 2D torus |
|---|
| All-reduce bandwidth | 340 TB/s per pod |
|---|
| Bisection bandwidth | 6.4 TB/s per pod |
|---|
| Power | 123 / 220 / 262 W (min/mean/max, measured) |
|---|
| Cooling | Liquid — a first for Google data centers |
|---|
Figures per Google Cloud’s current TPU v3 documentation 4. The 126-petaflops pod figure is Google’s current documented spec, which may differ from whatever number was stated live at the 2018 keynote; no official Google blog post confirming an exact keynote figure was found.
5.Where it’s made
SimpleStart here
Google does not disclose who manufactured TPU v3 or on what process. No official source names a foundry for this generation.
Deep diveThe technical detail
As with TPU v1 and v2, no official Google or Broadcom material ties a specific manufacturer to TPU v3. Google’s confirmed Broadcom manufacturing relationship is documented only from later generations.
6.Which systems use it
SimpleStart here
TPU v3 shipped as Cloud TPU Pods, rentable through Google Cloud, and remains available today as a legacy option 5.
Deep diveThe technical detail
Google’s own case studies from this generation include Recursion Pharmaceuticals, which cut a 24-hour local-GPU training run down to 15 minutes using TPU v3 Pods 6, and Google Translate, which used TPU v3 for both training and serving to push new models to production within hours of validation 7.
7.Official pricing
SimpleStart here
Google Cloud’s current pricing page lists TPU v3 as a legacy tier: $2.00/chip-hour for a pod allocation, $2.20/chip-hour for a standalone device 5.
Deep diveThe technical detail
Checked September 2026, from Google Cloud’s own pricing page 5: TPU v3 pod allocations at $2.00 per chip-hour, standalone devices at $2.20 per chip-hour in europe-west4. These are legacy rates, still available nearly eight years after launch.
9.What came before, what came next
SimpleStart here
Came before: TPU v2. Came after: TPU v4, announced May 2021, which added optical circuit switching to dynamically reconfigure the network between chips.
Deep diveThe technical detail
TPU v3 succeeded TPU v2 primarily on raw compute and cooling capacity rather than a new capability like training. The next generation, TPU v4, introduced a genuinely new architectural feature: optical circuit switches that let Google rewire the interconnect between chips on the fly 1.
10.Hidden in plain sight
SimpleStart here
A chip cooled by water quietly changed how fast Google Translate could ship new models — from a slow release cycle to hours.
Deep diveThe technical detail
Google’s own inference-records post states Google Translate could push new models into production “within hours” of validation once TPU v3 handled both training and serving 7 — a workflow change that traces directly back to the liquid-cooling breakthrough that let this generation run hot enough to be worth building at all.