AI Chips · Product page

Amazon Inferentia: The Complete Guide

The chip that started Amazon’s AI-silicon business — announced with almost no detail, then quietly filled in over the years that followed. Simple to expert, every number links to its source.

Stuck at any point? Ask an AI about this page →
Ask an AI about this chip:CAGPX
Amazon Inferentia 10 sections · 2 levels 10 linked sources Checked September 2026
How to read this page. Each section starts Simple, then goes to a Deep dive: stop wherever you have what you need. The small numbers are sources: click one to open the original document. Where the maker has not published a figure -- a price, a die size, a factory address -- this page says so rather than estimate it.

1.At a glance

SimpleStart here

Inferentia was Amazon’s first AI chip — built only to run already-trained AI models quickly and cheaply, not to train them. AWS announced it 28 November 2018, with the instances that actually use it, Inf1, arriving a year later 1 2.

Deep diveThe technical detail

At announcement, AWS gave almost no technical detail: then-AWS-CEO Andy Jassy described only that Inferentia chips would deliver “hundreds of TOPS (tera operations per second) of inference throughput” each, with “multiple AWS Inferentia chips” usable together for thousands of TOPS 1. AWS filled in the precise figures only later, on its own product page: “up to 128 tera operations per second (TOPS) of performance” per chip 3 — this page flags that the 2018 announcement and the current product page state two different levels of precision, both from AWS itself, rather than treating either alone as the full picture.

2.Launch and history

SimpleStart here

Amazon started its chip business with inference, not training — the cheaper, more common job of actually running an AI model once it already exists. Training-focused Trainium (this site’s sibling page) came two years later.

Deep diveThe technical detail

Inferentia was designed by Annapurna Labs, the Israeli chip company Amazon bought in 2015 4. Its first major public customer was Amazon’s own Alexa: by November 2020, AWS said the majority of Alexa’s workloads had moved to Inf1 instances, citing “25% lower end-to-end latency, and 30% lower cost compared to GPU-based instances for Alexa’s text-to-speech workloads” 5 — an early, concrete proof point for a chip that launched with almost no published detail.

3.What’s inside it

SimpleStart here

Each Inferentia chip has four NeuronCores — the processing units that do the actual inference math — plus 8GB of DDR4 memory and a separate pool of on-chip memory 3.

Deep diveThe technical detail

AWS’s own description: “The first-generation Inferentia has 8 GB of DDR4 memory per chip and also features a large amount of on-chip memory” 3 — AWS does not quantify that on-chip figure precisely. Data types supported: FP16, BF16 and INT8 3. AWS has never published a process node or power figure for Inferentia, and this page found no independent teardown or press report naming either — unlike later chips, where press has at least named a likely foundry (see the Trainium2 and Trainium3 pages on this site).

4.Full spec table

SimpleStart here

Up to 128 TOPS per chip, 8GB of DDR4 memory, four NeuronCores — sized for inference, not training.

Deep diveThe technical detail
SpecAmazon Inferentia
AI computeUp to 128 TOPS per chip 3
NeuronCores4 (first-generation) 3
Memory8 GB DDR4, plus unquantified on-chip memory 3
Data typesFP16, BF16, INT8 3
Process nodeNot published by AWS; no independent source found
Inf1 instancevCPUsMemoryChips
inf1.xlarge48 GiB1
inf1.2xlarge816 GiB1
inf1.6xlarge2448 GiB4
inf1.24xlarge96192 GiB16

Instance table per AWS’s own December 2019 launch post 2, unchanged on the current instance-types page 6.

5.Where it’s made

SimpleStart here

Designed by Annapurna Labs, Amazon’s in-house chip team. Amazon has never named the foundry that manufactures Inferentia 4.

Deep diveThe technical detail

Unlike Trainium2 and Trainium3, where press reports (not Amazon) name TSMC, this page found no report of any kind — press, analyst or teardown — naming a foundry for the original Inferentia. Coverage of the 2018 announcement noted explicitly that Amazon gave almost no technical detail at all: SiliconANGLE’s launch-day report says then-AWS-CEO Andy Jassy “provided few details on the design or specifications” 7, and Data Center Knowledge’s coverage the next day says he “did not share any details about the chip’s design or performance” 8. This page treats the foundry as genuinely unpublished rather than guessing from later Amazon chips.

6.Which systems use it

SimpleStart here

Rented through AWS as EC2 Inf1 instances, from a single chip up to 16 per server — never sold as a standalone chip 2.

Deep diveThe technical detail

Beyond Amazon’s own Alexa, AWS has published named customer results for Inf1: Sprinklr says it was “able to significantly improve the performance of one of our NLP models and improve the performance of one of our computer vision models” 9; Finch Computing reported inference costs down “over 80%” versus GPUs; Dataminr reported “up to 9x better throughput per dollar”; Autodesk reported “4.9x higher throughput” over a GPU instance for its NLU models; Condé Nast reported a “72% reduction in cost”; and Amazon’s own Search team reported inference costs down 85% 6. All are AWS-published customer quotes, not independently audited figures — Sprinklr’s own testimonial does not give a specific percentage, unlike the others.

7.Official pricing

SimpleStart here

Rented by the hour, not sold. AWS’s current published on-demand prices for Inf1 instances 6:

InstanceOn-demand1-yr reserved3-yr reserved
inf1.xlarge (1 chip)$0.228/hr$0.137/hr$0.101/hr
inf1.2xlarge (1 chip)$0.362/hr$0.217/hr$0.161/hr
inf1.6xlarge (4 chips)$1.180/hr$0.709/hr$0.525/hr
inf1.24xlarge (16 chips)$4.721/hr$2.835/hr$2.099/hr
Deep diveThe technical detail

Prices shown are AWS’s US East (default region) list prices as published on its Inf1 instance-types page; AWS shows a region selector and prices vary by region 6. There is no separate per-chip price — Inferentia has never been sold as a standalone part, only rented as part of an instance.

8.Real-world performance

SimpleStart here

AWS’s own comparison claims have changed over time. At the December 2019 launch: “up to 3x the inferencing throughput, and up to 40% lower cost per inference” versus AWS’s G4 (GPU) instances 2. On AWS’s current product page: “up to 2.3x higher throughput” and “up to 70% lower cost per inference” 6.

Deep diveThe technical detail

These are two different AWS-published figures from two different dates, not one consistent claim — this page states both rather than picking one, since AWS itself has not explained the change (a different GPU comparator, updated pricing, or simply a revised marketing figure are all plausible, unconfirmed explanations). Neither is an independent third-party benchmark; both are AWS’s own comparisons against its own GPU instances.

9.What came before, what came next

SimpleStart here
Came before: nothing — Inferentia was Amazon’s first AI chip. Came after: Inferentia2, covered on its own page on this site.
Deep diveThe technical detail

Amazon’s AI-silicon roadmap split into two families after this chip: Inferentia stayed inference-only through Inferentia2, while Trainium (announced 2020, covered on its own page on this site) took on training. In practice, Amazon’s newer Trainium chips now handle most inference too, according to the brand-level roadmap on this site 10 — Inferentia and Inferentia2 remain the dedicated inference-only line.

10.Hidden in plain sight

SimpleStart here

Amazon announced this chip with so little technical detail that press at the time explicitly noted the company would not describe its own design.

Deep diveThe technical detail

Contemporary coverage of the 28 November 2018 announcement is notably thin on Amazon’s own words: reports from the event describe then-AWS-CEO Andy Jassy giving “few details on the design or specifications” 7 and not sharing “any details about the chip’s design or performance” 8. The precise 128-TOPS figure, the 8GB DDR4 memory figure, and the four-NeuronCore count all appear only later, on AWS’s own product page, not in the launch announcement itself 1 3.

11.Sources

10 sources, checked September 2026. Where NVIDIA or another maker has not published a figure, this page says so rather than estimate it.

  1. AWS What’s New: Announcing Amazon InferentiaOfficial
  2. AWS News Blog: Amazon EC2 Update — Inf1 Instances with AWS Inferentia ChipsOfficial
  3. AWS: AI Chip — Amazon Inferentia (product page)Official
  4. Amazon Science: How silicon innovation became the secret sauce behind AWS’s successOfficial
  5. AWS News Blog: Majority of Alexa now running on faster, more cost-effective Amazon EC2 Inf1 instancesOfficial
  6. AWS: Amazon EC2 Inf1 instances pageOfficial
  7. SiliconANGLE (press): AWS launches its own custom chip for AI inferencing (28 Nov 2018)Press
  8. Data Center Knowledge (press): AWS unveils AI inference chip Inferentia (29 Nov 2018)Press
  9. AWS: Sprinklr case study — Amazon EC2 Inf1Official
  10. Clarigital: Amazon Trainium — brand guide (this site)Official

Ask an AI about this page

Opens your assistant with this page as the source, and a question rather than a summary. It will ask what you are building before it answers.

ChatGPTClaudeGeminiPerplexityGrok

Nothing is sent from here. The link carries only this page’s title and address.