Soaring demand for AI agents is creating a surprising computational crisis—not for GPUs, but for CPUs.
AWS executives called a meeting with engineers in May, delivering a stark warning: they must do everything possible to conserve computing resources to ensure their EC2 cloud server business can meet all future customer needs. This directive covers not only AI chips but also the traditional CPU servers that have long powered the internet.
This internal move reflects a broader shift: as more business employees use AI agents to develop software, CPU demand is skyrocketing, putting pressure on the entire computing infrastructure, not just GPU capacity. Jing Xie, co-founder and managing director of AI consulting firm Elendil Labs, noted that the per-person IT spending of his clients has doubled, directly due to the widespread use of AI agents.
Engineers now report that waiting times for CPU server resources have stretched from hours to days, a situation unprecedented in their careers. An engineer with several years at Amazon.com said this delay could affect project delivery timelines. Amazon.com responded that despite "high demand," it can still meet the computing needs of the "vast majority of internal and external customers" and is working closely with teams to use EC2 resources efficiently, as always. The company stated its guidelines for employees have not changed due to the current memory shortage.
External customers have been less affected so far. A consultant who helps businesses use AWS has not observed any operational shortages of contractually committed CPU capacity. However, Spot Instances—surplus server capacity sold at a steep discount but retrievable within two minutes—have become increasingly difficult to obtain in large quantities in recent months. This suggests that the gap between supply and demand for Amazon.com's overall computing power is narrowing, leaving less elasticity. For businesses relying on cheap spot instances for large workloads, this may signal rising cloud computing costs.
The core driver of this CPU shortage is the rapid adoption of agentic AI. Jing Xie explained that running AI agents typically requires more cloud resources, creating a sustained demand for CPUs. "Basically, a lot of the work being built, produced, and run now requires more CPU than it used to." Even within AI development, CPUs are essential for tasks like reading raw data from documents, images, and videos to prepare it for model training.
Chip makers have confirmed this trend. Intel CEO Lip-Bu Tan reported in April that the ratio of CPU to GPU usage in AI inference was 1:4. By July, Intel CFO David Zinsner said it had neared 1:1. Executives from AMD and Arm have made similar comments. Beyond the CPUs themselves, shortages of memory chips and limited physical space in data centers are also contributing to the crunch.
Balancing internal and external computing resources has become a common challenge for large tech companies. Google established a senior committee last year to coordinate computing allocation among its Cloud, DeepMind AI, and consumer businesses. Even so, internal friction remains, as demonstrated by the departure of prominent AI researcher Noam Shazeer over dissatisfaction with computing access. Microsoft's experience shows a different angle: better compute management can boost revenue, with CFO Amy Hood attributing some of Azure's growth to "efficiency gains" in managing CPU and GPU fleets.
Comments