AI Chips · Product page

Google TPU v5p: The Complete Guide

The chip behind Gemini, by Google’s own account. Its own two official pages disagree on one number by exactly 2x. Simple to expert, every number links to its source.

Ask an AI about this chip:CAGPX
Google TPU v5p 10 sections · 2 levels 6 linked sources Checked September 2026
How to read this page. Each section starts Simple, then goes to a Deep dive: stop wherever you have what you need. The small numbers are sources: click one to open the original document. Where the maker has not published a figure -- a price, a die size, a factory address -- this page says so rather than estimate it.

1.At a glance

SimpleStart here

TPU v5p is Google’s large-scale training chip, announced 6 December 2023. Google calls it “our most powerful, scalable, and flexible AI accelerator thus far,” and names it directly as the chip behind AI Hypercomputer, Google’s AI infrastructure system 1.

Deep diveThe technical detail

Google’s own launch post states TPU v5p is “tuned and orchestrated specifically for generative AI training and serving,” and in the same post says Gemini, Google’s flagship AI model family, was trained and is served using TPUs — without naming v5p as the specific chip, though the context makes v5p the natural read 1.

2.Launch and history

SimpleStart here

Where TPU v5e (announced three months earlier) optimized for cost, v5p optimized for raw scale — the largest pod Google had shipped at the time, aimed squarely at training the biggest AI models.

Deep diveThe technical detail

Google’s own claims versus TPU v4: more than double the FLOPS, 2.8 times faster training for large language models, 1.9 times faster for embedding-heavy models via second-generation SparseCores, triple the HBM capacity, and four times more scalable in total FLOPS available per pod 1.

3.What’s inside it

SimpleStart here

Two TensorCores and four second-generation SparseCores per chip 2 — a richer per-chip design than v5e’s single TensorCore, reflecting v5p’s focus on the largest training jobs rather than cost efficiency.

Deep diveThe technical detail

Each TensorCore carries four matrix-multiply units plus a vector and scalar unit 2. Pods are built from a base unit Google calls a “cube” — 64 chips across 16 hosts in a 4x4x4 arrangement — and only a full cube or larger gets genuine 3D-torus wraparound connectivity; smaller slices connect in 3D but without the wraparound 2. Google does not disclose HBM generation or process node for v5p.

4.Full spec table

SimpleStart here

459 teraflops per chip, 95GB of HBM memory, and pods scaling up to 8,960 chips 2.

Deep diveThe technical detail
SpecGoogle TPU v5p
Peak compute (bf16/fp8)459 TFLOPs per chip
Memory95 GiB HBM, 2,765 GBps bandwidth
TensorCores / SparseCores2 / 4 per chip
Interconnect topology3D torus
Pod size8,960 chips (built from 64-chip “cubes”)
Max single job6,144 chips (96 cubes)

Figures per Google Cloud’s TPU v5p documentation 2. Google’s own pages disagree on inter-chip interconnect bandwidth: the launch blog states 4,800 Gbps per chip, while the documentation states 1,200 GBps per chip — a 2x gap that may reflect a unidirectional-versus-bidirectional counting difference between the two pages, which this page reports rather than silently resolving 12.

5.Where it’s made

SimpleStart here

Google does not disclose who manufactures TPU v5p or on what process node. No official source names a foundry for this chip specifically.

Deep diveThe technical detail

As with v5e, Google’s Broadcom manufacturing relationship is documented at the company level for future TPU generations broadly, but no official filing or blog names a foundry tied to v5p by name.

6.Which systems use it

SimpleStart here

TPU v5p is available through Google Cloud in three regions: us-central1, us-east5 and europe-west4 3.

Deep diveThe technical detail

The accelerator type is ct5p-hightpu-4t in Google’s own tooling 2. Google’s v5p launch post says Gemini “was trained on, and is served, using TPUs” without naming v5p specifically by chip name — a reasonable inference given the post is v5p’s own launch announcement, but not a literal naming 1. Tracing the generations forward: Google’s release notes group v5p with v5e as the two chips Trillium (v6e) improves upon 4, and Trillium was in turn succeeded by Ironwood (TPU7x).

7.Official pricing

SimpleStart here

Google Cloud’s own pricing page lists TPU v5p on-demand at $4.20 per chip-hour in us-east5 and us-east1 5.

Deep diveThe technical detail

Checked September 2026, from Google Cloud’s own pricing page 5: on-demand $4.20/chip-hour in us-east5 (Columbus, Ohio) and us-east1 (South Carolina); a Dynamic Workload Scheduler option at $2.10/chip-hour (Flex-start) or $2.94/chip-hour (Calendar mode); committed-use discounts bring the rate down to $2.94/chip-hour (1-year) or $1.89/chip-hour (3-year).

8.Real-world performance

SimpleStart here

Google claims more than double the FLOPS of TPU v4, 2.8 times faster training for large language models, and 1.9 times faster embedding-heavy model training 1 — company-stated comparisons, not an independent benchmark.

Deep diveThe technical detail

This page reports Google’s own launch-announcement figures rather than a third-party MLPerf submission specific to v5p. The 1.9x embedding-model figure is attributed specifically to v5p’s second-generation SparseCores 1.

9.What came before, what came next

SimpleStart here

Came before: TPU v4. Came after: Trillium (TPU v6e), announced May 2024.

Deep diveThe technical detail

Google’s own release notes group v5e and v5p together as the two chips Trillium (v6e) improves upon 4. Google’s later Ironwood (TPU7x) announcement compares itself only to Trillium, not directly to v5p 6 — so the clean succession chain, per Google’s own material, runs v5p → Trillium (v6e) → Ironwood (TPU7x).

10.Hidden in plain sight

SimpleStart here

Two of Google’s own official pages for the same chip give a network-speed figure that differs by exactly double, and Google has not reconciled it.

Deep diveThe technical detail

TPU v5p’s launch blog states inter-chip interconnect bandwidth as 4,800 Gbps per chip 1 — 600 GBps once converted to the same unit. The current technical documentation for the same chip states 1,200 GBps per chip 2: exactly double the blog figure, not the same number restated. Both pages are official, current, and unambiguous in their own wording, and Google has not published anything reconciling the two, whether the difference reflects a unidirectional-versus-bidirectional counting convention or something else.

11.Sources

6 sources, checked September 2026. Where NVIDIA or another maker has not published a figure, this page says so rather than estimate it.

  1. Google Cloud Blog: Introducing Cloud TPU v5p and AI HypercomputerOfficial
  2. Google Cloud: TPU v5p DocumentationOfficial
  3. Google Cloud: Cloud TPU Regions and ZonesOfficial
  4. Google Cloud: Cloud TPU Release NotesOfficial
  5. Google Cloud: Cloud TPU PricingOfficial
  6. Google Blog: Ironwood: The First Google TPU for the Age of InferenceOfficial