How to read this page. Each section starts Simple, then goes to a Deep dive: stop wherever you have what you need. The small numbers are sources: click one to open the original document. Where the maker has not published a figure -- a price, a die size, a factory address -- this page says so rather than estimate it.
1.At a glance
SimpleStart here
Inferentia was Amazon’s first AI chip — built only to run already-trained AI models quickly and cheaply, not to train them. AWS announced it 28 November 2018, with the instances that actually use it, Inf1, arriving a year later 1 2.
Deep diveThe technical detail
At announcement, AWS gave almost no technical detail: then-AWS-CEO Andy Jassy described only that Inferentia chips would deliver “hundreds of TOPS (tera operations per second) of inference throughput” each, with “multiple AWS Inferentia chips” usable together for thousands of TOPS 1. AWS filled in the precise figures only later, on its own product page: “up to 128 tera operations per second (TOPS) of performance” per chip 3 — this page flags that the 2018 announcement and the current product page state two different levels of precision, both from AWS itself, rather than treating either alone as the full picture.
2.Launch and history
SimpleStart here
Amazon started its chip business with inference, not training — the cheaper, more common job of actually running an AI model once it already exists. Training-focused Trainium (this site’s sibling page) came two years later.
Deep diveThe technical detail
Inferentia was designed by Annapurna Labs, the Israeli chip company Amazon bought in 2015 4. Its first major public customer was Amazon’s own Alexa: by November 2020, AWS said the majority of Alexa’s workloads had moved to Inf1 instances, citing “25% lower end-to-end latency, and 30% lower cost compared to GPU-based instances for Alexa’s text-to-speech workloads” 5 — an early, concrete proof point for a chip that launched with almost no published detail.
3.What’s inside it
SimpleStart here
Each Inferentia chip has four NeuronCores — the processing units that do the actual inference math — plus 8GB of DDR4 memory and a separate pool of on-chip memory 3.
Deep diveThe technical detail
AWS’s own description: “The first-generation Inferentia has 8 GB of DDR4 memory per chip and also features a large amount of on-chip memory” 3 — AWS does not quantify that on-chip figure precisely. Data types supported: FP16, BF16 and INT8 3. AWS has never published a process node or power figure for Inferentia, and this page found no independent teardown or press report naming either — unlike later chips, where press has at least named a likely foundry (see the Trainium2 and Trainium3 pages on this site).
4.Full spec table
SimpleStart here
Up to 128 TOPS per chip, 8GB of DDR4 memory, four NeuronCores — sized for inference, not training.
Deep diveThe technical detail
| Spec | Amazon Inferentia |
|---|
| AI compute | Up to 128 TOPS per chip 3 |
|---|
| NeuronCores | 4 (first-generation) 3 |
|---|
| Memory | 8 GB DDR4, plus unquantified on-chip memory 3 |
|---|
| Data types | FP16, BF16, INT8 3 |
|---|
| Process node | Not published by AWS; no independent source found |
|---|
| Inf1 instance | vCPUs | Memory | Chips |
|---|
| inf1.xlarge | 4 | 8 GiB | 1 |
| inf1.2xlarge | 8 | 16 GiB | 1 |
| inf1.6xlarge | 24 | 48 GiB | 4 |
| inf1.24xlarge | 96 | 192 GiB | 16 |
Instance table per AWS’s own December 2019 launch post 2, unchanged on the current instance-types page 6.
5.Where it’s made
SimpleStart here
Designed by Annapurna Labs, Amazon’s in-house chip team. Amazon has never named the foundry that manufactures Inferentia 4.
Deep diveThe technical detail
Unlike Trainium2 and Trainium3, where press reports (not Amazon) name TSMC, this page found no report of any kind — press, analyst or teardown — naming a foundry for the original Inferentia. Coverage of the 2018 announcement noted explicitly that Amazon gave almost no technical detail at all: SiliconANGLE’s launch-day report says then-AWS-CEO Andy Jassy “provided few details on the design or specifications” 7, and Data Center Knowledge’s coverage the next day says he “did not share any details about the chip’s design or performance” 8. This page treats the foundry as genuinely unpublished rather than guessing from later Amazon chips.
6.Which systems use it
SimpleStart here
Rented through AWS as EC2 Inf1 instances, from a single chip up to 16 per server — never sold as a standalone chip 2.
Deep diveThe technical detail
Beyond Amazon’s own Alexa, AWS has published named customer results for Inf1: Sprinklr says it was “able to significantly improve the performance of one of our NLP models and improve the performance of one of our computer vision models” 9; Finch Computing reported inference costs down “over 80%” versus GPUs; Dataminr reported “up to 9x better throughput per dollar”; Autodesk reported “4.9x higher throughput” over a GPU instance for its NLU models; Condé Nast reported a “72% reduction in cost”; and Amazon’s own Search team reported inference costs down 85% 6. All are AWS-published customer quotes, not independently audited figures — Sprinklr’s own testimonial does not give a specific percentage, unlike the others.
7.Official pricing
SimpleStart here
Rented by the hour, not sold. AWS’s current published on-demand prices for Inf1 instances 6:
| Instance | On-demand | 1-yr reserved | 3-yr reserved |
|---|
| inf1.xlarge (1 chip) | $0.228/hr | $0.137/hr | $0.101/hr |
| inf1.2xlarge (1 chip) | $0.362/hr | $0.217/hr | $0.161/hr |
| inf1.6xlarge (4 chips) | $1.180/hr | $0.709/hr | $0.525/hr |
| inf1.24xlarge (16 chips) | $4.721/hr | $2.835/hr | $2.099/hr |
Deep diveThe technical detail
Prices shown are AWS’s US East (default region) list prices as published on its Inf1 instance-types page; AWS shows a region selector and prices vary by region 6. There is no separate per-chip price — Inferentia has never been sold as a standalone part, only rented as part of an instance.
9.What came before, what came next
SimpleStart here
Came before: nothing — Inferentia was Amazon’s first AI chip.
Came after: Inferentia2, covered on its own page on this site.
Deep diveThe technical detail
Amazon’s AI-silicon roadmap split into two families after this chip: Inferentia stayed inference-only through Inferentia2, while Trainium (announced 2020, covered on its own page on this site) took on training. In practice, Amazon’s newer Trainium chips now handle most inference too, according to the brand-level roadmap on this site 10 — Inferentia and Inferentia2 remain the dedicated inference-only line.
10.Hidden in plain sight
SimpleStart here
Amazon announced this chip with so little technical detail that press at the time explicitly noted the company would not describe its own design.
Deep diveThe technical detail
Contemporary coverage of the 28 November 2018 announcement is notably thin on Amazon’s own words: reports from the event describe then-AWS-CEO Andy Jassy giving “few details on the design or specifications” 7 and not sharing “any details about the chip’s design or performance” 8. The precise 128-TOPS figure, the 8GB DDR4 memory figure, and the four-NeuronCore count all appear only later, on AWS’s own product page, not in the launch announcement itself 1 3.