AWS and NVIDIA Plan 2 Million More GPUs for Agentic and Physical AI
AWS and NVIDIA announced plans to deploy 2 million additional NVIDIA GPUs across AWS infrastructure in 2027–2028 while expanding integration across CPUs, networking, open models, accelerated data systems and robotics.
A planned expansion far beyond a single GPU purchase
AWS and NVIDIA announced on August 26, 2026 that they plan to deploy 2 million additional NVIDIA GPUs across AWS global infrastructure during 2027–2028. The announcement is important not only because of the scale of accelerator capacity, but because the companies are extending their collaboration across the surrounding AI stack: CPUs, networking, cloud security, open models, data processing, vector indexing and robotics.
The language matters: this is a forward deployment plan, not a statement that two million new GPUs are already online. Capacity will arrive over time, and the practical effect for customers will depend on region, instance availability, pricing, networking, power and the maturity of the software integrations described in the announcement.
Agentic AI is pushing infrastructure beyond raw accelerator counts
The expansion includes Blackwell Ultra, Rubin and Rubin Ultra systems as well as work to bring NVIDIA Vera CPU-based infrastructure to AWS. For agentic applications, this reflects a broader architecture problem. Large-scale agents can combine model inference with retrieval, tool calls, data pipelines, code execution, memory and repeated reasoning loops. Their bottleneck is therefore not always the model accelerator alone. CPU throughput, storage, networking, vector search and isolation become part of end-to-end performance.
AWS and NVIDIA are consequently deepening integration with the AWS Nitro System and Elastic Fabric Adapter while also extending accelerated data processing and vector indexing. The strategy suggests that cloud providers expect a growing share of AI workloads to look like distributed applications rather than isolated model endpoints.
Open models and data infrastructure are part of the same platform play
NVIDIA's Nemotron family remains available through Amazon Bedrock and SageMaker, giving AWS customers another managed and self-deployment path for NVIDIA-backed open models. The announcement also describes GPU acceleration for data preparation on Amazon EMR and vector-index construction on Amazon OpenSearch.
That combination matters because production AI systems often spend significant time outside the model itself. Feature engineering, ETL, document processing and index building can determine latency and cost before a prompt ever reaches an LLM. Bringing more of those stages onto accelerated infrastructure can reduce the gap between model benchmarks and application-level performance, although the actual economics will vary substantially by workload.
Physical AI broadens the target market
The partnership also reaches into robotics. Amazon Robotics is using NVIDIA's physical-AI platform across simulation, synthetic-data generation, robot training and validation. This makes the infrastructure announcement relevant to warehouses and industrial systems as well as conventional generative-AI services.
Physical AI has a different deployment profile from cloud-only software: training and simulation may be centralized, while inference and control can move closer to machines. The AWS–NVIDIA roadmap therefore links large cloud training capacity with tools intended to support real-world robotic systems.
What to watch next
The most useful signals will be operational rather than headline GPU counts: when the new instance families become broadly available, which regions receive capacity, how pricing changes, whether Vera-based systems improve agent workloads, how quickly Rubin generations enter production, and whether accelerated OpenSearch and EMR integrations materially reduce total application cost.
For developers and enterprises, the announcement reinforces a clear trend: the competitive unit in AI infrastructure is increasingly the full production stack—accelerators, CPUs, networking, security, models, data systems and deployment tooling—rather than any single chip generation.
This article is built from the source material below. Open the originals for full context and the latest updates.