How to read this page. Each section starts Simple, then goes to a Deep dive: stop wherever you have what you need. The small numbers are sources: click one to open the original document. Where the maker has not published a figure -- a price, a die size, a factory address -- this page says so rather than estimate it.
1.At a glance
SimpleStart here
Trillium is Google’s sixth-generation TPU, announced 15 May 2024 at Google I/O and generally available from 16 December 2024 12. Google positions it as a general-purpose workhorse for training and serving, without the inference-only framing it would later give the next generation.
Deep diveThe technical detail
Google’s own numbers versus TPU v5e: 4.7 times the peak compute per chip, double the HBM capacity and bandwidth, double the inter-chip interconnect bandwidth, and 67% better energy efficiency 1.
2.Launch and history
SimpleStart here
Trillium is Google’s answer to a generation of increasingly long-context, multimodal AI models — a chip built to handle bigger context windows and more varied data types, not one narrow workload.
Deep diveThe technical detail
Google’s launch blog names Essential AI, Nuro, Deep Genomics, Deloitte, Google DeepMind (for Gemini training and serving) and Lightricks as early adopters 1. Google shipped Trillium roughly seven months after the announcement, reaching general availability on 16 December 2024 2.
3.What’s inside it
SimpleStart here
One TensorCore per chip with a wider matrix-multiply unit and higher clock speed than the prior generation, plus a third-generation SparseCore for recommendation-style models 1.
Deep diveThe technical detail
Google’s documentation discloses BF16 (918 TFLOPs) and INT8 (1,836 TOPs) figures for Trillium, but publishes no FP8 figure at all for this generation 3 — unlike the following generation, Ironwood, whose documentation leads with an FP8 number. Google does not disclose a process node or foundry for Trillium.
4.Full spec table
SimpleStart here
918 teraflops (bf16) per chip, 32GB of HBM memory, connected up to 256 chips per pod 3.
Deep diveThe technical detail
| Spec | Google Trillium (TPU v6e) |
|---|
| Peak compute (bf16) | 918 TFLOPs per chip |
|---|
| Peak compute (int8) | 1,836 TOPs per chip |
|---|
| Memory | 32 GB HBM, 1,638 GBps bandwidth |
|---|
| Inter-chip interconnect | 800 GBps bidirectional, 4 ports, 2D torus |
|---|
| Pod size | 256 chips |
|---|
| Pod peak (bf16) | 234.9 PFLOPs |
|---|
| Chips per host | 8 |
|---|
Figures per Google Cloud’s TPU v6e documentation 3.
5.Where it’s made
SimpleStart here
Google does not disclose who manufactures Trillium or on what process node. No official Google or Broadcom source names a foundry for this chip specifically.
Deep diveThe technical detail
Broadcom’s April 2026 SEC filing on its long-term TPU agreement with Google discusses “future generations” broadly, without naming Trillium, Ironwood, or any other generation by name, and does not disclose a foundry such as TSMC. This page does not assert one.
6.Which systems use it
SimpleStart here
Trillium is available through Google Cloud across five regions spanning North America, Europe, Asia-Pacific and South America 4.
Deep diveThe technical detail
Regions include us-central1, us-east1, us-east5, europe-west4, asia-northeast1 and southamerica-west1 4. Google’s carbon-efficiency reporting shows Trillium delivered a 20% reduction in Compute Carbon Intensity compared to the prior fleet period, reaching 125 grams of CO2-equivalent per exaflop, based on Google’s own fleet data through January 2026 5.
7.Official pricing
SimpleStart here
Google Cloud’s own pricing page lists Trillium on-demand from $2.70 per chip-hour in the US, rising to $3.24 in Tokyo 6.
Deep diveThe technical detail
| Region | On-demand ($/chip-hr) | 3-year committed use |
|---|
| us-east1, us-east5 | $2.70 | $1.22 |
| europe-west4 | $2.97 | ~$1.35 |
| asia-northeast1 | $3.24 | $1.46 |
Figures per Google Cloud’s own pricing page, checked September 2026 6. A Dynamic Workload Scheduler option is also listed at $1.35/chip-hour (Flex-start) or $1.89/chip-hour (Calendar mode) across regions; Google does not display a page-level “last updated” date or a fixed spot-pricing figure.
9.What came before, what came next
SimpleStart here
Came before: TPU v5e and TPU v5p. Came after: Ironwood (TPU7x), announced April 2025, Google’s first TPU built specifically for inference rather than general use.
Deep diveThe technical detail
Google’s own release notes describe Trillium’s improvements “over the prior generations, v5e and v5p” together 2. Ironwood succeeded it roughly eleven months later, with Google explicitly framing that next chip around “the age of inference” rather than Trillium’s general-purpose positioning.
10.Hidden in plain sight
SimpleStart here
Google discloses two precision levels for Trillium’s speed — and quietly skips the one every AI headline currently cares about most.
Deep diveThe technical detail
Trillium’s official documentation states BF16 and INT8 performance figures, but no FP8 number appears anywhere in Google’s published specs for this generation 3 — notable because FP8 became the precision format most AI chip vendors led with by 2024–2025 for inference workloads. The very next generation, Ironwood, reverses this: its documentation leads with an FP8 figure as the headline compute number.