AI Chips · Product page

Google TPU v3: The Complete Guide

So much heat that Google plumbed its data centers for the first time. Then it processed 1 million images a second. Simple to expert, every number links to its source.

Ask an AI about this chip:CAGPX
Google TPU v3 10 sections · 2 levels 7 linked sources Checked September 2026
How to read this page. Each section starts Simple, then goes to a Deep dive: stop wherever you have what you need. The small numbers are sources: click one to open the original document. Where the maker has not published a figure -- a price, a die size, a factory address -- this page says so rather than estimate it.

1.At a glance

SimpleStart here

TPU v3 ran so hot that Google plumbed liquid cooling into its data centers for the first time — a first Google’s own retrospective calls out by name 1. Announced and shown entering alpha around Google I/O and Google Cloud Next ’18.

Deep diveThe technical detail

Google Cloud Next ’18 (24–26 July 2018) confirmed “Cloud TPU Pods and TPU v3 are now available in alpha” 2. The exact figure Google gave live on stage at I/O for peak pod performance is not independently confirmed in an official Google blog post; this page uses Google’s current documented spec instead of the press-reported keynote number.

2.Launch and history

SimpleStart here

TPU v3 pushed enough compute into each chip that air cooling could no longer keep up. Google’s own infrastructure team says liquid cooling let it double chip density and quadruple the size of its AI supercomputer compared to the air-cooled generation before it.

Deep diveThe technical detail

Google’s own words: “Google first used liquid cooling in TPU v3 that was deployed in 2018,” and that switch let Google “double chip density” and build a liquid-cooled supercomputer “four times” the size of its air-cooled TPU v2 predecessor 3. This is confirmed at two separate official Google Cloud sources 13.

3.What’s inside it

SimpleStart here

TPU v3 moved to HBM2 memory and a 2D torus network connecting up to 1,024 chips into one pod 4 — a step up in both memory technology and scale from TPU v2.

Deep diveThe technical detail

Per Google’s own current documentation 4: 32 GiB of HBM2 per chip at 900 GB/s, a 2D torus interconnect topology, 340 TB/s of pod-wide all-reduce bandwidth and 6.4 TB/s of bisection bandwidth. Google does not publish a process node for TPU v3.

4.Full spec table

SimpleStart here

123 teraflops per chip, 32GB of HBM2 memory, and a 1,024-chip pod delivering 126 petaflops 4.

Deep diveThe technical detail
SpecGoogle TPU v3
Peak compute (per chip, bf16)123 teraflops
Peak compute (per pod, bf16)126 petaflops
Memory32 GiB HBM2 per chip, 900 GB/s
Pod size1,024 chips, 2D torus
All-reduce bandwidth340 TB/s per pod
Bisection bandwidth6.4 TB/s per pod
Power123 / 220 / 262 W (min/mean/max, measured)
CoolingLiquid — a first for Google data centers

Figures per Google Cloud’s current TPU v3 documentation 4. The 126-petaflops pod figure is Google’s current documented spec, which may differ from whatever number was stated live at the 2018 keynote; no official Google blog post confirming an exact keynote figure was found.

5.Where it’s made

SimpleStart here

Google does not disclose who manufactured TPU v3 or on what process. No official source names a foundry for this generation.

Deep diveThe technical detail

As with TPU v1 and v2, no official Google or Broadcom material ties a specific manufacturer to TPU v3. Google’s confirmed Broadcom manufacturing relationship is documented only from later generations.

6.Which systems use it

SimpleStart here

TPU v3 shipped as Cloud TPU Pods, rentable through Google Cloud, and remains available today as a legacy option 5.

Deep diveThe technical detail

Google’s own case studies from this generation include Recursion Pharmaceuticals, which cut a 24-hour local-GPU training run down to 15 minutes using TPU v3 Pods 6, and Google Translate, which used TPU v3 for both training and serving to push new models to production within hours of validation 7.

7.Official pricing

SimpleStart here

Google Cloud’s current pricing page lists TPU v3 as a legacy tier: $2.00/chip-hour for a pod allocation, $2.20/chip-hour for a standalone device 5.

Deep diveThe technical detail

Checked September 2026, from Google Cloud’s own pricing page 5: TPU v3 pod allocations at $2.00 per chip-hour, standalone devices at $2.20 per chip-hour in europe-west4. These are legacy rates, still available nearly eight years after launch.

8.Real-world performance

SimpleStart here

Google reports 32 TPU v3 devices processed 1 million images a second in an official MLPerf Inference benchmark, and TPU v3 Pods ran over 84% faster than the fastest on-premises systems in MLPerf Training 76.

Deep diveThe technical detail

Google’s own November 2019 post states 32 Cloud TPU v3 devices processed 1 million ResNet-50 images per second in MLPerf Inference v0.5, with near-linear scaling from 1 to 32 devices, and calculated that processing all 7.7 billion of the world’s photographs at that rate would cost under $600 and take under 2.5 hours 7. A separate July 2019 post reports TPU v3 Pods running over 84% faster than the fastest on-premises systems measured in MLPerf Training 6. Both are independently audited MLPerf results, not purely company-stated comparisons.

9.What came before, what came next

SimpleStart here

Came before: TPU v2. Came after: TPU v4, announced May 2021, which added optical circuit switching to dynamically reconfigure the network between chips.

Deep diveThe technical detail

TPU v3 succeeded TPU v2 primarily on raw compute and cooling capacity rather than a new capability like training. The next generation, TPU v4, introduced a genuinely new architectural feature: optical circuit switches that let Google rewire the interconnect between chips on the fly 1.

10.Hidden in plain sight

SimpleStart here

A chip cooled by water quietly changed how fast Google Translate could ship new models — from a slow release cycle to hours.

Deep diveThe technical detail

Google’s own inference-records post states Google Translate could push new models into production “within hours” of validation once TPU v3 handled both training and serving 7 — a workflow change that traces directly back to the liquid-cooling breakthrough that let this generation run hot enough to be worth building at all.

11.Sources

7 sources, checked September 2026. Where NVIDIA or another maker has not published a figure, this page says so rather than estimate it.

  1. Google Cloud: Ten Years of TPUs: A Look BackOfficial
  2. Google Cloud Blog: What Happened at Google Cloud Next ’18Official
  3. Google Cloud Blog: Enabling 1MW IT Racks and Liquid Cooling at OCP EMEA SummitOfficial
  4. Google Cloud: TPU v3 DocumentationOfficial
  5. Google Cloud: Cloud TPU PricingOfficial
  6. Google Cloud Blog: Cloud TPU Pods Break AI Training RecordsOfficial
  7. Google Cloud Blog: Cloud TPU Breaks Scalability Records for AI InferenceOfficial