AWS (Amazon Web Services) and Nvidia are expanding their long-running collaboration as demand for AI infrastructure continues to grow. The companies plan to deploy 2 million additional Nvidia GPUs across AWS data centres worldwide in 2027 and 2028. This will expand cloud capacity for businesses, research organisations and governments running large-scale AI workloads.
The expanded collaboration will also cover CPUs, networking, AI factories, open models and robotics. AWS said the agreement will give customers more flexibility to build and run demanding AI workloads across its cloud infrastructure.
The additional capacity will support workloads including agentic AI, scientific research, enterprise automation and physical AI.
The announcement follows AWS's earlier plan to add more than 1 million Nvidia GPUs by 2026. According to AWS, demand for AI computing capacity has been stronger than initially expected.
“Customers want the freedom to choose the best tools for their AI workloads,” AWS CEO Matt Garman said, highlighting the company's focus on flexible AI infrastructure.
Nvidia CEO Jensen Huang said the expanded partnership will cover the broader AI computing ecosystem, including GPUs, CPUs, networking, software and open models.
AWS is also working to bring NVIDIA Vera CPU-based infrastructure to its cloud platform. The CPUs are designed to handle workloads associated with agentic AI and reinforcement learning.
These workloads can include code execution, data processing, tool use, analytics and AI-agent orchestration. Vera CPUs can be paired with Nvidia GPUs or used as standalone computing resources.
The firms are also beefing up their efforts on Nvidia NVLink Fusion. AWS announced that its Annapurna Labs team will collaborate with Nvidia on custom high-bandwidth memory technology for future Trainium chips. The aim is to enhance the memory performance and power efficiency while enabling Trainium and Nvidia GPUs to operate on a shared rack-scale architecture.
AWS and Nvidia will build AI data centres for the US government, with up to 100,000 Nvidia GPUs on secure AWS infrastructure. The systems will be able to support federal and national security workloads that require Impact Level 6 (IL6) security or higher.
The partnership also covers existing AI services. Nemotron, Amazon Bedrock, and Amazon SageMaker enable Nvidia to offer open models to customers. According to AWS, Nvidia cuDF's GPU-accelerated data processing can offer up to 3.7 times faster processing for Apache Spark than a CPU-based configuration and up to 30% improved price-performance.
Amazon Robotics leverages Nvidia's Jetson, Omniverse, and Isaac platforms to create simulations, synthetic data, and real-world testing to advance warehouse robotics automation. The latest pledges come on the heels of almost 16 years of cooperation between AWS and Nvidia.
Pragya is a Technology reporter with over four years of experience in digital media and content writing. She holds a Master’s degree in Journalism and has covered a wide range of stories spanning space, smartphones, gadgets, artificial intelligence, emerging technologies and the ways technology is transforming everyday life. She focuses on breaking down complex technology and science developments into clear, engaging, and reader-friendly stories, with a keen interest in emerging trends and their real-world impact. Before joining her current newsroom, Pragya worked with News9Live, where she covered the technology beat extensively, reporting on smartphones, consumer technology, AI, space and science, while also contributing to video and visual content. Her experience includes breaking news, explainers, SEO-driven stories, product coverage, interviews, unboxing videos and live event reporting. She has also covered major technology and AI events, giving her experience in both newsroom and on-ground reporting. Beyond journalism, Pragya is an avid gamer and a passionate reader of fiction. She enjoys exploring immersive worlds through games and books, with a particular interest in stories that offer new perspectives, ideas, and experiences.
Catch all the Business News, Market News, Breaking News Events and Latest News Updates on Live Mint. Download The Mint News App to get Daily Market Updates.
Oops! Looks like you have exceeded the limit to bookmark the image. Remove some to bookmark this image.
Facts Only
* AWS and Nvidia plan to deploy 2 million additional Nvidia GPUs across AWS data centers in 2027 and 2028.
* The expanded collaboration covers CPUs, networking, AI factories, open models, and robotics.
* The additional capacity will support agentic AI, scientific research, enterprise automation, and physical AI workloads.
* AWS is working to bring NVIDIA Vera CPU-based infrastructure to its cloud platform for agentic AI and reinforcement learning workloads.
* NVLink Fusion collaboration involves AWS Annapurna Labs working with Nvidia on custom high-bandwidth memory technology for future Trainium chips.
* AI data centers will be built for the US government, supporting up to 100,000 Nvidia GPUs on secure AWS infrastructure requiring Impact Level 6 (IL6) security or higher.
* Nemotron, Amazon Bedrock, and Amazon SageMaker enable Nvidia to offer open models to customers.
* Nvidia cuDF offers processing speeds up to 3.7 times faster for Apache Spark than a CPU-based configuration and up to 30% improved price-performance compared to CPU configurations.
* Amazon Robotics leverages Nvidia's Jetson, Omniverse, and Isaac platforms.
Executive Summary
AWS and Nvidia are expanding their collaboration to deploy 2 million additional Nvidia GPUs across AWS data centers globally in 2027 and 2028. This expansion aims to increase cloud capacity for businesses, research organizations, and governments running large-scale AI workloads. The expanded partnership encompasses CPUs, networking, AI factories, open models, and robotics, providing customers with greater flexibility for building and running demanding AI workloads on AWS infrastructure.
The increased capacity will support diverse workloads, including agentic AI, scientific research, enterprise automation, and physical AI. AWS is also developing Vera CPU-based infrastructure to handle agentic AI and reinforcement learning tasks, which can be paired with Nvidia GPUs or used independently for functions like code execution and data processing. Furthermore, the collaboration includes advancements in NVLink Fusion for memory technology between Trainium chips and Nvidia GPUs, and the establishment of secure AI data centers for the US government with up to 100,000 Nvidia GPUs. Existing services like Amazon Bedrock and SageMaker facilitate access to open models, and Nvidia's cuDF offers performance improvements for data processing over CPU configurations.
Full Take
The expansion of the AWS-Nvidia partnership signals a structural shift where foundational AI infrastructure is becoming deeply integrated and commoditized across hardware, software, and application layers. The focus on expanding capacity beyond just GPUs to encompass CPUs, networking, and specialized memory solutions like Vera CPUs indicates a move toward building holistic, end-to-end AI stacks rather than optimizing discrete components. This integration supports the emergence of agentic AI, which requires not only massive parallel processing but also complex orchestration capabilities (code execution, tool use, orchestration), suggesting that future value will reside less in raw compute and more in system-level efficiency and control.
The strategic focus on joint memory technology (NVLink Fusion for Trainium) suggests a pattern of internalizing performance bottlenecks to ensure monolithic system coherence across the stack. The government data center initiative implies a separation between commercial deployment flexibility and high-security, sovereign infrastructure capability. This dynamic creates tension: while the partnership promises unparalleled flexibility for customers, it establishes Nvidia as the central hardware linchpin upon which AWS constructs its future AI offering. The integration of robotics platforms further solidifies an ecosystem where simulation and physical execution are inseparable from the underlying cloud architecture. The unstated assumption is that this level of deep co-development will effectively define the next generation of enterprise AI paradigms, meaning the patterns being established today will dictate systemic constraints in the near future regarding proprietary vs. open models, and centralized security protocols.
What assumptions about technological leadership and infrastructure control underpin this coordinated expansion? What costs are implicitly shifted to end-users or governments by prioritizing ecosystem integration over pure, unconstrained hardware specification? Does this structure create an attractive single point of failure for future AI deployment if the symbiotic relationship were to fracture? What independent metrics exist outside of capability expansion that define the true long-term success of such deep technical alignment?
