How to read this page. Each section starts Simple, then goes to a Deep dive: stop wherever you have what you need. The small numbers are sources: click one to open the original document. Where the maker has not published a figure -- a price, a die size, a factory address -- this page says so rather than estimate it.
1.At a glance
SimpleStart here
Inferentia2 is AWS’s second dedicated inference chip, built to run the much larger generative AI models that arrived after the original Inferentia — chatbots, image generators, and other models with tens of billions of parameters. Inf2 instances became generally available 13 April 2023 1.
Deep diveThe technical detail
AWS positioned Inferentia2 specifically around scale: the chip adds high-bandwidth memory (HBM) and supports models distributed across multiple chips — a capability the original, DDR4-memory Inferentia was not built for (see the Inferentia page on this site). AWS extended Inf2 to more regions through 2023, including Mumbai in December 2.
2.Launch and history
SimpleStart here
Inferentia2 arrived as generative AI took off — its job was making Amazon’s inference line capable of running large language and image models, not just the smaller classic-ML models the original Inferentia targeted.
Deep diveThe technical detail
AWS’s own documentation groups Inferentia2’s architecture together with the original Trainium chip, describing them as sharing a common NeuronCore design 3 — the two chips launched around the same period and share the same underlying compute core, even though one is inference-only and the other trains models too (see the Trainium page on this site).
3.What’s inside it
SimpleStart here
AWS documents Inferentia2’s architecture together with the original Trainium chip — both use the same NeuronCore design, built around a large matrix-multiply engine 3.
Deep diveThe technical detail
Per AWS’s own Neuron architecture documentation, the shared NeuronCore design pairs a systolic-array tensor engine for matrix math with a vector engine, a scalar engine, and a programmable “GpSimd” engine for custom code 3 — the same four-engine layout this site’s Trainium page describes in more detail. AWS has not published a separate, Inferentia2-specific architecture diagram distinct from this shared documentation.
4.Full spec table
SimpleStart here
Up to 190 TFLOPS at FP16, 32GB of HBM memory per chip — enough to hold and run much larger models than the original Inferentia’s 8GB of DDR4.
Deep diveThe technical detail
| Spec | Amazon Inferentia2 |
|---|
| AI compute | Up to 190 TFLOPS FP16 4 |
|---|
| Memory | 32 GB HBM 4 |
|---|
| Process node | Not published by AWS 4 |
|---|
AWS’s own product page gives these as the only chip-level figures for Inferentia2; more detailed figures (memory bandwidth, per-core throughput) are not published the way they are for Trainium2 and Trainium3 (see those pages on this site) 4.
5.Where it’s made
SimpleStart here
AWS has not named a foundry for Inferentia2. Unlike Trainium2 and Trainium3, this page found no press report naming one either.
Deep diveThe technical detail
Press coverage of Amazon’s chip supply chain (cited on the Trainium2 and Trainium3 pages on this site) discusses TSMC by name for those two chips specifically, but this page found no equivalent reporting for Inferentia2 — its foundry appears to be genuinely unreported, not simply omitted from this page.
6.Which systems use it
SimpleStart here
Rented through AWS as EC2 Inf2 instances, up to 12 chips per server, available across multiple AWS regions including Mumbai since December 2023 2.
Deep diveThe technical detail
AWS has not published a customer-quotes section for Inf2 as detailed as the one on its Inf1 page (see the Inferentia page on this site) or its Trainium2/Trainium3 pages. This page found no equivalent set of named, quoted Inferentia2 customers in AWS’s own published material.
7.Official pricing
SimpleStart here
Rented by the hour, not sold. AWS’s current published on-demand prices for Inf2 instances 5:
| Instance | On-demand | 1-yr reserved | 3-yr reserved |
|---|
| inf2.xlarge (1 chip) | $0.76/hr | $0.45/hr | $0.30/hr |
| inf2.48xlarge (12 chips) | $12.98/hr | $7.79/hr | $5.19/hr |
Deep diveThe technical detail
As with every Amazon inference chip on this site, there is no separate per-chip sale price — only a rental rate for the instance it comes packaged in 5. AWS has not published a Capacity Block (advance-booking) price for Inferentia2, unlike some Trainium generations (see the Trainium and Trainium2 pages on this site).
9.What came before, what came next
SimpleStart here
Came before: the original Inferentia, covered on its own page on this site.
Came after: no third Inferentia generation has been announced; Amazon’s newer Trainium chips now handle a growing share of inference workloads instead.
Deep diveThe technical detail
Amazon has not announced an Inferentia3. Its brand-level roadmap states that Trainium2 “powers the majority of inference on Amazon Bedrock,” Amazon’s own AI-model service 6 — suggesting Amazon’s inference roadmap has shifted toward Trainium rather than a dedicated third Inferentia generation, though Amazon has not said so directly.
10.Hidden in plain sight
SimpleStart here
Amazon publishes noticeably less detail about Inferentia2 than about its Trainium chips — no customer-quote page, no Capacity Block price, no independently reported foundry.
Deep diveThe technical detail
Compared with Trainium2 and Trainium3, each of which has a dedicated customer-testimonial page, named press-reported foundry, and (for Trainium2) an advance-booking Capacity Block price, Inferentia2’s public documentation is comparatively thin — consistent with Amazon’s own roadmap emphasis shifting toward Trainium for newer inference workloads 6.