How to read this page. Each section starts Simple, then goes to a Deep dive: stop wherever you have what you need. The small numbers are sources: click one to open the original document. Where the maker has not published a figure -- a price, a die size, a factory address -- this page says so rather than estimate it.
1.At a glance
SimpleStart here
TPU v5p is Google’s large-scale training chip, announced 6 December 2023. Google calls it “our most powerful, scalable, and flexible AI accelerator thus far,” and names it directly as the chip behind AI Hypercomputer, Google’s AI infrastructure system 1.
Deep diveThe technical detail
Google’s own launch post states TPU v5p is “tuned and orchestrated specifically for generative AI training and serving,” and in the same post says Gemini, Google’s flagship AI model family, was trained and is served using TPUs — without naming v5p as the specific chip, though the context makes v5p the natural read 1.
2.Launch and history
SimpleStart here
Where TPU v5e (announced three months earlier) optimized for cost, v5p optimized for raw scale — the largest pod Google had shipped at the time, aimed squarely at training the biggest AI models.
Deep diveThe technical detail
Google’s own claims versus TPU v4: more than double the FLOPS, 2.8 times faster training for large language models, 1.9 times faster for embedding-heavy models via second-generation SparseCores, triple the HBM capacity, and four times more scalable in total FLOPS available per pod 1.
3.What’s inside it
SimpleStart here
Two TensorCores and four second-generation SparseCores per chip 2 — a richer per-chip design than v5e’s single TensorCore, reflecting v5p’s focus on the largest training jobs rather than cost efficiency.
Deep diveThe technical detail
Each TensorCore carries four matrix-multiply units plus a vector and scalar unit 2. Pods are built from a base unit Google calls a “cube” — 64 chips across 16 hosts in a 4x4x4 arrangement — and only a full cube or larger gets genuine 3D-torus wraparound connectivity; smaller slices connect in 3D but without the wraparound 2. Google does not disclose HBM generation or process node for v5p.
4.Full spec table
SimpleStart here
459 teraflops per chip, 95GB of HBM memory, and pods scaling up to 8,960 chips 2.
Deep diveThe technical detail
| Spec | Google TPU v5p |
|---|
| Peak compute (bf16/fp8) | 459 TFLOPs per chip |
|---|
| Memory | 95 GiB HBM, 2,765 GBps bandwidth |
|---|
| TensorCores / SparseCores | 2 / 4 per chip |
|---|
| Interconnect topology | 3D torus |
|---|
| Pod size | 8,960 chips (built from 64-chip “cubes”) |
|---|
| Max single job | 6,144 chips (96 cubes) |
|---|
Figures per Google Cloud’s TPU v5p documentation 2. Google’s own pages disagree on inter-chip interconnect bandwidth: the launch blog states 4,800 Gbps per chip, while the documentation states 1,200 GBps per chip — a 2x gap that may reflect a unidirectional-versus-bidirectional counting difference between the two pages, which this page reports rather than silently resolving 12.
5.Where it’s made
SimpleStart here
Google does not disclose who manufactures TPU v5p or on what process node. No official source names a foundry for this chip specifically.
Deep diveThe technical detail
As with v5e, Google’s Broadcom manufacturing relationship is documented at the company level for future TPU generations broadly, but no official filing or blog names a foundry tied to v5p by name.
6.Which systems use it
SimpleStart here
TPU v5p is available through Google Cloud in three regions: us-central1, us-east5 and europe-west4 3.
Deep diveThe technical detail
The accelerator type is ct5p-hightpu-4t in Google’s own tooling 2. Google’s v5p launch post says Gemini “was trained on, and is served, using TPUs” without naming v5p specifically by chip name — a reasonable inference given the post is v5p’s own launch announcement, but not a literal naming 1. Tracing the generations forward: Google’s release notes group v5p with v5e as the two chips Trillium (v6e) improves upon 4, and Trillium was in turn succeeded by Ironwood (TPU7x).
7.Official pricing
SimpleStart here
Google Cloud’s own pricing page lists TPU v5p on-demand at $4.20 per chip-hour in us-east5 and us-east1 5.
Deep diveThe technical detail
Checked September 2026, from Google Cloud’s own pricing page 5: on-demand $4.20/chip-hour in us-east5 (Columbus, Ohio) and us-east1 (South Carolina); a Dynamic Workload Scheduler option at $2.10/chip-hour (Flex-start) or $2.94/chip-hour (Calendar mode); committed-use discounts bring the rate down to $2.94/chip-hour (1-year) or $1.89/chip-hour (3-year).
9.What came before, what came next
SimpleStart here
Came before: TPU v4. Came after: Trillium (TPU v6e), announced May 2024.
Deep diveThe technical detail
Google’s own release notes group v5e and v5p together as the two chips Trillium (v6e) improves upon 4. Google’s later Ironwood (TPU7x) announcement compares itself only to Trillium, not directly to v5p 6 — so the clean succession chain, per Google’s own material, runs v5p → Trillium (v6e) → Ironwood (TPU7x).
10.Hidden in plain sight
SimpleStart here
Two of Google’s own official pages for the same chip give a network-speed figure that differs by exactly double, and Google has not reconciled it.
Deep diveThe technical detail
TPU v5p’s launch blog states inter-chip interconnect bandwidth as 4,800 Gbps per chip 1 — 600 GBps once converted to the same unit. The current technical documentation for the same chip states 1,200 GBps per chip 2: exactly double the blog figure, not the same number restated. Both pages are official, current, and unambiguous in their own wording, and Google has not published anything reconciling the two, whether the difference reflects a unidirectional-versus-bidirectional counting convention or something else.