AWS Custom Silicon History and Timeline - Nitro System, Graviton, Inferentia, Trainium, and Introduction
First Published:
Last Updated:
The individual Amazon EC2 instance families that carry these chips are covered in the companion Amazon EC2 Instance Types History and Timeline, and the broader service history in AWS History and Timeline regarding Amazon EC2. This article zooms in on the silicon itself — when each chip family and generation was announced and became generally available — so that questions such as "when did AWS release Graviton?", "what is the AWS Nitro System?", or "what is the difference between Inferentia and Trainium?" can be answered from a single page.
This article focuses on the chip-level milestones: the announcement and the general availability of each generation. It intentionally does not duplicate the per-instance-size and per-Region rollout detail that the instance types timeline already covers, and it does not include pricing or independent performance comparisons.
Background and Method of Creating AWS Custom Silicon Historical Timeline
AWS custom silicon began with the 2015 acquisition of Annapurna Labs, a microelectronics company that AWS had already been working with on the first hardware-offload components of what would become the AWS Nitro System. From that single team, four distinct silicon families emerged: Nitro (hardware offload and the lightweight hypervisor), Graviton (general-purpose Arm CPUs), Inferentia (machine learning inference), and Trainium (machine learning training).The networking-offload building blocks that preceded the Nitro name are part of the same lineage — Enhanced Networking with SR-IOV first appeared on the C3 family in 2013, and the Elastic Network Adapter (ENA) followed in 2016 — and the first instance built entirely on the Nitro System, the C5 family, arrived in 2017. Every current-generation Amazon EC2 instance, whether x86, Arm, or accelerator-based, now runs on Nitro.
- Tracking the history of AWS-designed chips and organizing the transition of generations across the Nitro, Graviton, Inferentia, and Trainium families
- Summarizing the current four-family silicon lineup, how the families map to Amazon EC2 instances, and the role of the AWS Neuron SDK
A note on scope (important): each row in the timeline below marks a chip-family milestone — the announcement or the general availability of a chip generation. Announcements (often at AWS re:Invent) and general availability are recorded as separate rows because they can be months, or even more than a year, apart. Individual instance sizes, per-Region rollouts, and minor derivative variants are not given their own rows; those belong to the companion instance types timeline.
The content posted is limited to major milestones related to the current AWS custom silicon families and necessary for the family overview.
In other words, please note that the items on this timeline are not all AWS silicon updates, but are representative milestones that I have picked out.
There may be slight variations in dates due to differences between the What's New posting date and the AWS News Blog posting date for the same milestone.

AWS Custom Silicon Historical Timeline (Updates from 2015)
Here is a timeline of AWS custom silicon milestones, from the 2015 Annapurna Labs acquisition to the present. Announcements and general availability (GA) are listed as separate rows.2015 | 2017 | 2018 | 2019 | 2020 | 2021 | 2022 | 2023 | 2024 | 2025 | 2026
* The table can be sorted by clicking on the column names.| Date | Summary |
|---|---|
| 2015 | AWS acquires Annapurna Labs, the chip team behind its custom silicon. AWS references the 2015 acquisition in its own materials but has not published an exact announcement date; Annapurna Labs had already been working with AWS on the first hardware-offload components that became the AWS Nitro System. |
| 2017-11-06 | Amazon EC2 C5 becomes generally available as the first instance family built on the AWS Nitro System. Nitro moves virtualization, networking, and storage onto dedicated hardware and introduces a lightweight hypervisor, replacing the earlier Xen-based stack. |
| 2017-11-28 | AWS formally introduces the name "AWS Nitro System" at AWS re:Invent 2017, with EC2 Bare Metal instances in preview. The Nitro System is described as a collection of AWS-built hardware offload and security components; the dedicated Nitro Cards and the Nitro Security Chip are what later make bare metal and AWS Graviton instances possible. |
| 2018-11-26 | AWS announces its first-generation AWS Graviton processor with the Amazon EC2 A1 family at AWS re:Invent 2018. Graviton is a custom 64-bit Arm CPU; A1 is the first Arm-based EC2 instance and runs on the AWS Nitro System. |
| 2018-11-28 | AWS announces AWS Inferentia, its first purpose-built machine learning inference chip, at AWS re:Invent 2018. Inferentia is designed to run trained models in production and is programmed through the AWS Neuron SDK. |
| 2019-12-03 | AWS previews AWS Graviton2 with the upcoming M6g, C6g, and R6g instances at AWS re:Invent 2019. Graviton2 is a second-generation Arm design that spans the general-purpose, compute-optimized, and memory-optimized instance categories. |
| 2019-12-03 | Amazon EC2 Inf1 instances, powered by AWS Inferentia, become generally available. Inf1 is the first production vehicle for Inferentia and targets inference workloads such as image recognition, recommendation, and natural language processing. |
| 2020-05-11 | Amazon EC2 M6g, the first AWS Graviton2 instance family, becomes generally available. M6g is the general-purpose member of the Graviton2 lineup, followed by the compute- and memory-optimized siblings. |
| 2020-06-11 | Amazon EC2 C6g and R6g, the compute- and memory-optimized AWS Graviton2 families, become generally available. Together with M6g they complete the first full Graviton2 generation across the three core instance categories. |
| 2020-10-28 | AWS Nitro Enclaves becomes generally available. Built on the Nitro Hypervisor and previewed at re:Invent 2019, Nitro Enclaves creates isolated compute environments for sensitive data, with no persistent storage, interactive access, or external networking. |
| 2020-12-01 | AWS announces AWS Trainium, a chip purpose-built for machine learning training, at AWS re:Invent 2020. Trainium is the training counterpart to Inferentia and shares the AWS Neuron SDK and support for major deep learning frameworks. |
| 2021-11-30 | AWS previews AWS Graviton3 with the Amazon EC2 C7g instance at AWS re:Invent 2021. Graviton3 is the third-generation Arm design and first appears in the compute-optimized C7g family. |
| 2021-11-30 | AWS opens the preview of Amazon EC2 Trn1 instances, the first vehicle for AWS Trainium, at AWS re:Invent 2021. Trn1 is aimed at training large deep learning models, including large language models. |
| 2022-05-11 | NitroTPM and UEFI Secure Boot become generally available for Amazon EC2, extending the AWS Nitro System. NitroTPM provides a virtual Trusted Platform Module for storing secrets and boot measurements, and UEFI Secure Boot validates the boot chain. |
| 2022-05-23 | Amazon EC2 C7g, the first AWS Graviton3 instance family, becomes generally available. C7g brings Graviton3 to compute-optimized workloads and, like all Graviton families, runs on the Nitro System. |
| 2022-10-10 | Amazon EC2 Trn1 instances, powered by AWS Trainium, become generally available. Trn1 is the first generally available Trainium vehicle and uses Elastic Fabric Adapter (EFA) networking to scale training across many instances. |
| 2022-11-28 | AWS previews AWS Graviton3E with the upcoming C7gn and Hpc7g instances at AWS re:Invent 2022. Graviton3E is a variant of Graviton3 tuned for high-performance computing and network-intensive workloads. |
| 2022-11-29 | AWS opens the preview of Amazon EC2 Inf2 instances, powered by AWS Inferentia2, at AWS re:Invent 2022. Inferentia2 is the second-generation inference chip, designed for large models including generative AI. |
| 2023-04-13 | Amazon EC2 Inf2 instances, powered by AWS Inferentia2, become generally available. Inf2 targets inference for large language and diffusion models and supports distributed inference across accelerators. |
| 2023-04-13 | Amazon EC2 Trn1n instances, a network-optimized AWS Trainium variant, become generally available. Trn1n increases network bandwidth over Trn1 for tightly coupled, large-scale training clusters. |
| 2023-06-20 | Amazon EC2 Hpc7g and C7gn, powered by AWS Graviton3E, become generally available. Hpc7g targets tightly coupled HPC and C7gn targets network-intensive workloads, both using Graviton3E. |
| 2023-11-28 | AWS previews AWS Graviton4 with the Amazon EC2 R8g instance at AWS re:Invent 2023. Graviton4 is the fourth-generation Arm design and first appears in the memory-optimized R8g family. |
| 2023-11-28 | AWS announces AWS Trainium2, its next-generation machine learning training chip, at AWS re:Invent 2023. Trainium2 is designed for training and deploying the largest foundation models and later ships in the Trn2 instances and UltraServers. |
| 2024-07-09 | Amazon EC2 R8g, the first AWS Graviton4 instance family, becomes generally available. R8g brings Graviton4 to memory-optimized workloads such as databases and in-memory analytics. |
| 2024-12-03 | Amazon EC2 Trn2 instances become generally available and Trn2 UltraServers enter preview at AWS re:Invent 2024. Trn2, powered by AWS Trainium2, targets large-scale generative AI training and inference; UltraServers connect multiple Trn2 instances into a single high-bandwidth compute node. |
| 2025-12-02 | Amazon EC2 Trn3 UltraServers, powered by AWS Trainium3, become generally available at AWS re:Invent 2025. Trainium3 is the third-generation training chip, delivered first as UltraServers for the largest foundation-model workloads. |
| 2025-12-04 | AWS previews AWS Graviton5 with the Amazon EC2 M9g instance at AWS re:Invent 2025. Graviton5 is the fifth-generation Arm design and first appears in the general-purpose M9g family in preview. |
| 2026-06-10 | Amazon EC2 M9g and M9gd, the first AWS Graviton5 instances, become generally available. M9g is the general-purpose member of the Graviton5 generation, with M9gd adding local NVMe instance storage. AWS describes both as built on the sixth-generation AWS Nitro System and as the first instances to feature the Nitro Isolation Engine. |
| 2026-06-11 | The AWS Nitro Isolation Engine becomes generally available on all AWS Graviton5-based instances. First announced at AWS re:Invent 2025, the Nitro Isolation Engine is a purpose-built component inside the Nitro Hypervisor that mediates all access to virtual machine memory, CPU register state, and I/O devices. AWS applies formal verification to it and describes the result as the first formally verified cloud hypervisor. |
| 2026-06-30 | Amazon EC2 C9g and C9gd, the compute-optimized AWS Graviton5 instances, become generally available. C9g extends Graviton5 to compute-optimized workloads. |
Current Overview, Functions, Features of AWS Custom Silicon
As of 2026, AWS custom silicon is organized into four families, each with a distinct purpose. The section below describes what each family is for, how the families map to Amazon EC2 instances, their use cases, and the AWS Neuron SDK that programs the machine learning accelerators.The Four AWS Silicon Families
- AWS Nitro System (foundation): the hardware and lightweight-hypervisor platform that underlies all current-generation Amazon EC2 instances. Nitro offloads virtualization, VPC networking, and EBS and instance storage to dedicated Nitro Cards, enforces a hardware root of trust with the Nitro Security Chip, and uses the lightweight Nitro Hypervisor. Nitro is what makes bare metal instances, AWS Graviton instances, Nitro Enclaves, and hardware-accelerated networking possible. The platform has itself advanced through generations: the AWS Graviton5 instances that shipped in 2026 are described by AWS as built on the sixth-generation Nitro System and are the first to carry the Nitro Isolation Engine, a formally verified component of the Nitro Hypervisor that mediates all access to instance memory, CPU register state, and I/O devices.
- AWS Graviton (general-purpose CPU): AWS-designed 64-bit Arm processors that power the general-purpose, compute-optimized, and memory-optimized instances carrying the
gsuffix. Graviton has advanced through five generations — Graviton1 (A1, 2018), Graviton2 (2019–2020), Graviton3 and Graviton3E (2021–2023), Graviton4 (2023–2024), and Graviton5 (M9g preview 2025, GA 2026). - AWS Inferentia (ML inference): accelerators purpose-built for running trained machine learning models in production. Inferentia (Inf1) targets classic deep learning inference, and Inferentia2 (Inf2) targets large models including generative AI.
- AWS Trainium (ML training): accelerators purpose-built for training deep learning models. Trainium (Trn1/Trn1n), Trainium2 (Trn2 and Trn2 UltraServers), and Trainium3 (Trn3 UltraServers) scale from single instances to tightly coupled clusters for the largest foundation models, and are also used for inference.
How AWS Silicon Maps to Amazon EC2 Instances
The four families surface to customers primarily as Amazon EC2 instances. The table below shows how each silicon family maps to the instances that carry it; the per-instance detail is covered in the companion Amazon EC2 Instance Types History and Timeline.| Silicon family | Purpose | Representative Amazon EC2 instances |
|---|---|---|
| AWS Nitro System | Foundation: hardware offload and lightweight hypervisor | All current-generation instances (C5 onward), plus bare metal and Nitro Enclaves |
| AWS Graviton | General-purpose Arm CPU | Instances with the g suffix, such as A1, T4g, M6g–M9g, C7g, R8g, plus Hpc7g and C7gn (Graviton3E) |
| AWS Inferentia | Machine learning inference | Inf1 (Inferentia), Inf2 (Inferentia2) |
| AWS Trainium | Machine learning training (and inference) | Trn1, Trn1n, Trn2 and Trn2 UltraServers, Trn3 UltraServers |
AWS Custom Silicon Use Cases
- Running general-purpose workloads — web and application servers, microservices, and small-to-medium databases — on AWS Graviton for price-performance.
- Migrating x86 workloads to Arm-based AWS Graviton where application and binary compatibility allow.
- Serving machine learning inference (recommendation, search, vision, and language models) on AWS Inferentia.
- Training and fine-tuning deep learning models, including large language models, on AWS Trainium.
- Running high-performance computing and network-intensive workloads on AWS Graviton3E (Hpc7g, C7gn).
- Processing highly sensitive data in isolated environments using AWS Nitro Enclaves.
- Meeting security and boot-integrity requirements with the Nitro Security Chip, NitroTPM, and UEFI Secure Boot.
Specific Examples of Use Cases
- Re-platforming a fleet of x86 web servers onto M-series Graviton instances (for example, M6g through M9g) to improve price-performance.
- Hosting an in-memory database or cache on memory-optimized Graviton (R6g, R8g) or on Graviton-based instances generally.
- Deploying a computer-vision or recommendation inference endpoint on Inf1, then moving a large generative model to Inf2.
- Pre-training or fine-tuning a large language model on a cluster of Trn1n or Trn2 instances connected with Elastic Fabric Adapter, or on Trn2/Trn3 UltraServers for the largest models.
- Running tightly coupled weather, genomics, or CFD simulations on Hpc7g.
- Isolating cryptographic key material or personally identifiable information inside a Nitro Enclave attached to a parent EC2 instance.
AWS Custom Silicon Key Functions and Features
- Hardware offload: the Nitro Cards move networking, storage, and virtualization off the main CPU so that nearly all host resources are available to customer workloads.
- Lightweight hypervisor: the Nitro Hypervisor provides isolation with minimal overhead and enables bare metal instances. On AWS Graviton5-based instances it includes the Nitro Isolation Engine, the component AWS formally verifies to prove that isolation between instances and from AWS operators holds.
- Hardware root of trust: the Nitro Security Chip, NitroTPM, and UEFI Secure Boot protect and validate the boot chain and platform integrity.
- Arm CPU generations: AWS Graviton advances roughly every generation across general-purpose, compute-optimized, and memory-optimized families, with an
Evariant (Graviton3E) for HPC and networking. - Purpose-built ML acceleration: AWS Inferentia targets inference and AWS Trainium targets training, both exposing their accelerators through the AWS Neuron SDK.
- Scale-out training and inference: Elastic Fabric Adapter networking and UltraServers connect many accelerators into tightly coupled clusters.
- Confidential and isolated compute: AWS Nitro Enclaves provides isolated environments with no persistent storage, interactive access, or external networking.
The AWS Neuron SDK
AWS Inferentia and AWS Trainium are programmed through the AWS Neuron SDK, a single software stack that compiles, runs, and profiles models on both chip families. Neuron integrates with popular machine learning frameworks so that models built with them can target Inferentia and Trainium accelerators. Because the two families share Neuron, teams can adopt a common toolchain for both training on Trainium and inference on Inferentia.References:
AWS Nitro System
AWS Graviton Processor
AWS Trainium
AWS Inferentia
Frequently Asked Questions about AWS Custom Silicon History
When did AWS acquire Annapurna Labs?
AWS acquired Annapurna Labs in 2015. AWS references the year in its own materials — for example, the AWS News Blog post announcing the Arm-based A1 instances states that AWS acquired Annapurna Labs in 2015 after working with the team on the first version of the AWS Nitro System — but it has not published an exact acquisition date. Annapurna Labs is the team behind all four AWS custom silicon families: Nitro, Graviton, Inferentia, and Trainium.What is the AWS Nitro System?
The AWS Nitro System is the hardware and lightweight-hypervisor platform underlying all current-generation Amazon EC2 instances. It debuted with the C5 family, which became generally available on November 6, 2017, and was formally named the "AWS Nitro System" at AWS re:Invent 2017 (November 28, 2017). Nitro offloads virtualization, VPC networking, and EBS and instance storage to dedicated Nitro Cards, enforces a hardware root of trust with the Nitro Security Chip, and uses the lightweight Nitro Hypervisor, which is what makes bare metal instances, AWS Graviton instances, and Nitro Enclaves possible. The platform has advanced through generations of its own: AWS describes the Graviton5-based instances released in 2026 as built on the sixth-generation Nitro System and as the first to include the Nitro Isolation Engine, a formally verified component of the Nitro Hypervisor.When did AWS release Graviton, and how many generations are there?
AWS announced its first Graviton processor with the Amazon EC2 A1 family at AWS re:Invent 2018 (November 26, 2018). There are five generations to date: Graviton1 (A1, 2018); Graviton2, previewed at re:Invent 2019 with M6g reaching general availability on May 11, 2020; Graviton3, previewed at re:Invent 2021 with C7g generally available on May 23, 2022, plus Graviton3E for HPC and networking (Hpc7g and C7gn, 2023); Graviton4, previewed at re:Invent 2023 with R8g generally available on July 9, 2024; and Graviton5, previewed with the M9g family at re:Invent 2025 (December 4, 2025), with M9g/M9gd reaching general availability on June 10, 2026 and the compute-optimized C9g/C9gd following on June 30, 2026.What is the difference between AWS Inferentia and AWS Trainium?
AWS Inferentia is purpose-built for machine learning inference — running already-trained models in production — and appears in the Inf1 (Inferentia, generally available December 3, 2019) and Inf2 (Inferentia2, generally available April 13, 2023) instances. AWS Trainium is purpose-built for training deep learning models and appears in Trn1 (generally available October 10, 2022), Trn2 and Trn2 UltraServers (generally available December 3, 2024), and Trn3 UltraServers (Trainium3, generally available December 2, 2025); Trainium is also used for inference. Both families are programmed through the shared AWS Neuron SDK.What is the AWS Neuron SDK?
The AWS Neuron SDK is the software stack that compiles, runs, and profiles machine learning models on AWS Inferentia and AWS Trainium. It integrates with popular deep learning frameworks and provides a common toolchain across both chip families, so the same stack is used for training on Trainium and inference on Inferentia.What is the difference between AWS Graviton and x86 instances?
AWS Graviton instances use AWS-designed 64-bit Arm processors and carry theg suffix (for example m7g, c7g, r8g), while x86 instances use Intel processors (the i suffix, for example m7i) or AMD processors (the a suffix, for example m7a). All three run on the AWS Nitro System. Choosing between them is primarily a matter of application and binary compatibility (Arm versus x86) and price-performance for the specific workload. The instance-level detail is covered in the companion Amazon EC2 Instance Types timeline.All entries in the timeline above are linked to their primary AWS source (a What's New announcement, an AWS News Blog post, or an official Amazon announcement). This article is reviewed regularly to incorporate new AWS custom silicon milestones.
Summary
In this article, I gathered the history of AWS custom silicon into a single timeline — from the 2015 acquisition of Annapurna Labs through the four families that grew out of it: the AWS Nitro System that underpins virtualization, the AWS Graviton general-purpose Arm CPUs, and the AWS Inferentia and AWS Trainium machine learning accelerators, up to the Graviton5, Trainium3, and Inferentia2 generations of the mid-2020s.Read alongside the instance-level Amazon EC2 Instance Types History and Timeline, which follows how each instance family and generation carrying these chips was introduced, and the service-level AWS History and Timeline regarding Amazon EC2. For where this silicon runs in practice, see how the Nitro System underpins isolated execution environments in How AWS Lambda Execution Environments Work, and how AWS Inferentia and AWS Trainium are used for model serving in Self-Managed LLM Inference on Amazon EKS.
In addition, there is also a historical timeline of all AWS services including services beyond custom silicon, so please have a look if you are interested.
AWS History and Timeline - Almost All AWS Services List, Announcements, General Availability(GA)
This timeline will be updated as AWS custom silicon continues to evolve.
References:
AWS Nitro System
AWS Compute Blog (AWS Nitro Isolation Engine: Formally verifying the hypervisor in the AWS Nitro System)
AWS Graviton Processor
AWS Trainium
AWS Inferentia
AWS Documentation (Amazon EC2 instance type naming conventions)
What's New with AWS?
AWS News Blog
References:
Tech Blog with curated related content
Written by Hidekazu Konishi