How to read this page. Each section starts Simple, then goes to a Deep dive: stop wherever you have what you need. The small numbers are sources: click one to open the original document. Where the maker has not published a figure -- a price, a die size, a factory address -- this page says so rather than estimate it.
1.At a glance
SimpleStart here
Trainium3 is AWS’s newest, most powerful training chip and its first built on a 3-nanometre process. AWS first announced it 3 December 2024, the same day Trainium2 shipped 1, and Trn3 UltraServers went on sale a year later, 2 December 2025 2.
Deep diveThe technical detail
AWS calls Trainium3 “the first 3nm AI chip from AWS” 3. By April 2026, CEO Andy Jassy wrote that Trainium3 had “just started shipping at the start of 2026,” was 30 to 40% more price-performant than Trainium2, and was already “nearly fully-subscribed” 4 — a demand signal similar to Trainium2’s own sellout (see that page on this site), reached within months rather than years.
2.Launch and history
SimpleStart here
Trainium3 kept the fast, roughly annual cadence Amazon established with Trainium2 — announced the same day its predecessor shipped, and on sale about a year after that.
Deep diveThe technical detail
A TechCrunch tour of Amazon’s Trainium lab in Austin, Texas (March 2026) described Trainium3 as liquid-cooled and detailed the lab’s modest scale — “about the size of two large conference rooms,” with a separate test data centre in a co-location facility 5. That is press description of Amazon’s facilities, not an AWS-published detail.
3.What’s inside it
SimpleStart here
Trainium3’s cores support new, faster 8-bit and 4-bit number formats (MXFP8 and MXFP4) at four times the speed of the BF16 format Trainium2 relied on, plus hardware support for routing in “mixture of experts” models 6.
Deep diveThe technical detail
AWS’s Neuron documentation names the new core design NeuronCore-v4: alongside MXFP8/MXFP4 at 4× BF16 speed, it speeds up the exponential math used in attention layers fourfold 6. Its UltraServer uses a new interconnect, NeuronSwitch-v1, an all-to-all fabric with twice the bandwidth within each UltraServer compared with Trainium2’s equivalent 3. Notably, Trainium3’s BF16 speed per core did not increase over Trainium2’s — both run at 79 TFLOPS BF16 per core; the entire generational gain comes from the new 8-bit and 4-bit formats, not from faster traditional math 7 6.
4.Full spec table
SimpleStart here
2.52 PFLOPS of FP8 compute per chip, 144GB of HBM3e memory, up to 144 chips in one UltraServer — AWS’s own figures.
Deep diveThe technical detail
| Spec | Amazon Trainium3 |
|---|
| AI compute (FP8) | 2.52 PFLOPS per chip 2 |
|---|
| Memory | 144 GB HBM3e 2 |
|---|
| Memory bandwidth | 4.9 TB/s (AWS product page) or 4.7 TB/s (AWS Neuron docs) |
|---|
| Process | 3-nanometre 2 |
|---|
| System | Chips | AI compute (FP8) | Memory |
|---|
| Trn3 UltraServer | Up to 144 Trainium3 | 362 PFLOPS | 20.7 TB |
AWS says a Trn3 UltraServer has up to 4.4× the performance and 4× the performance per watt of an equivalent Trainium2 UltraServer — AWS’s own claim, not an independent benchmark 2. AWS does not state whether the 2.52-petaflop chip figure is dense or sparse; its Neuron documentation separately lists 315 teraflops of MXFP8 per core across eight cores (315 × 8 = 2.52 PFLOPS), suggesting the published figure is dense 6.
5.Where it’s made
SimpleStart here
AWS has not named a foundry for Trainium3. Press reports name TSMC, on a 3-nanometre process 8.
Deep diveThe technical detail
This is a press attribution, not an AWS statement — TrendForce reported, before the chip even shipped, that Trainium3 was “reportedly built with TSMC’s 3nm” process 8. AWS’s own materials confirm the 3-nanometre process node directly 2 3 but do not name the manufacturer — consistent with Amazon’s pattern across every chip on this site of never naming a foundry itself.
6.Which systems use it
SimpleStart here
AWS names Anthropic, Karakuri, Metagenomi, NetoAI, Ricoh, Splash Music and Decart as named Trainium3 customers, several reporting up to 50% lower costs 3.
Deep diveThe technical detail
Trn3 UltraServers link through EC2 UltraClusters 3.0, which AWS says can connect up to a million Trainium chips in one cluster 3 — the infrastructure scale behind Project Rainier (see the Trainium2 page on this site) extended to the newer generation.
7.Official pricing
SimpleStart here
Not published. Unlike Trainium and Trainium2, AWS has not published an on-demand price or a Capacity Block price for Trainium3 as of this page’s check date 9 10.
Deep diveThe technical detail
This is a genuine gap, not an oversight on this page’s part: AWS’s own Capacity Block pricing page lists rates for Trainium and Trainium2 but no Trainium3 entry, and its Trn3 instance-types page carries no on-demand price table either. For comparison, see this site’s Trainium and Trainium2 pages, both of which do have published AWS prices.
9.What came before, what came next
SimpleStart here
Came before: Trainium2, covered on its own page on this site.
Came after: Trainium4 — announced but, as of this page’s September 2026 check date, still roughly 18 months from broad availability by Amazon’s own account, and not given its own page on this site.
Deep diveThe technical detail
Amazon’s own most recent public language on Trainium4, from CEO Andy Jassy’s April 2026 shareholder letter: “A significant chunk of Trainium4, which is still about 18 months from broad availability, has already been reserved” 4 — explicitly pre-production. Trainium4 was previewed, not launched, at AWS re:Invent in December 2025, with AWS describing it only in future terms: at least six times Trainium3’s FP4 performance, three times its FP8 performance, and four times its memory bandwidth 3. No Meta-MTIA-style ambiguous press report exists for Trainium4 the way it did for Meta’s MTIA 400 (see this site’s Meta MTIA 300 page) — Amazon’s own timeline is the only public source, and it is unambiguous: not yet shipped. Trainium4 will support NVIDIA’s NVLink Fusion connection, a first for an Amazon-designed chip, but AWS has not yet said who manufactures it or on what process 3.
10.Hidden in plain sight
SimpleStart here
Trainium3’s single biggest generational gain — the leap to 8-bit and 4-bit math — comes with no improvement at all to its traditional 16-bit compute speed.
Deep diveThe technical detail
AWS’s own Neuron documentation shows Trainium3’s BF16 throughput per core is identical to Trainium2’s: 79 TFLOPS either way 7 6. Every headline compute gain — the 2.52 PFLOPS FP8 figure, the 4.4× UltraServer performance claim — comes from new lower-precision number formats (MXFP8, MXFP4) rather than a faster core. A workload that cannot use those lower-precision formats would see little benefit from upgrading, a nuance AWS’s own headline comparisons do not spell out.