Artificial intelligence (AI) agents are proliferating at an explosive pace, triggering an unexpected computing crisis. This time, however, the shortage is not in graphics processing units (GPUs), but in the central processing units (CPUs) that have long underpinned the internet era. Market intelligence indicates that senior management at Amazon Web Services (AWS) recently issued strict directives to internal engineers to conserve computing power. Some engineers report that wait times for CPU server resources have stretched from a few hours to several days, signaling that the pressure of computing scarcity has fully spread to traditional general-purpose computing infrastructure.
According to a report by tech publication The Information, AWS management convened engineers for a meeting in May to deliver a cautionary signal: to ensure that its core EC2 cloud server business can meet all customer demand in the future, engineers must do everything possible to conserve computing resources. This directive covers not only the chronically supply-constrained AI-specific chips but also explicitly targets traditional CPU servers.
Sources familiar with the matter revealed that AWS has set deadlines for various teams later this year, requiring them to reduce their computing resource footprint. Engineers are progressively shutting down EC2 virtual server instances that were previously used for software development but now sit idle, aiming to reallocate this computing power to external paying customers. An engineer who has worked at AWS for several years disclosed that obtaining CPU server resources, a process that once took only a few hours, now requires queuing for days—a situation unprecedented in their career. They bluntly stated that such delays are highly likely to impede project delivery timelines.
In response to external concerns, an AWS spokesperson emphasized that the company can still meet the computing needs of the vast majority of its internal and external customers despite massive demand. The company stated it is working closely with internal teams to ensure EC2 resources are utilized as efficiently as ever. AWS also clarified that the resource usage guidelines provided to employees have not changed due to the current tight supply of memory chips.
While the impact on external customers remains relatively limited for now, warning signs have emerged in AWS’s “Spot Instances” market. Spot Instances, which are surplus server resources sold at steep discounts but can be reclaimed with just two minutes’ notice, have become increasingly difficult to secure in large batches in recent months. Analysts suggest this indicates that the gap between AWS’s overall computing supply and demand is narrowing, compressing its flexibility buffer. For enterprise customers reliant on low-cost Spot Instances to run large-scale workloads, this may foreshadow impending upward pressure on cloud computing costs.
Agentic AI Proliferation Upends Computing Demand Structure
The core driver behind this CPU shortage is the rapid proliferation of Agentic AI. Jing Xie, Co-founder and Managing Director of AI consultancy Elendil Labs, analyzed that AI agents typically require calling upon more cloud computing resources during execution, thereby generating sustained demand for CPUs. She observed that per-capita IT spending among clients has doubled, directly attributable to the large-scale use of AI agents.
“Essentially, a lot of the work being built, produced, and run now requires more CPUs than before,” Jing Xie stated. Even within the AI development pipeline itself, CPUs play an indispensable role. For instance, when AI companies prepare data for model training, they need CPUs to read raw data from files, images, and videos.
Chip Giants Confirm CPU-to-GPU Usage Nears 1:1
Statements from major chipmakers further corroborate this trend. Intel CEO Lip-Bu Tan disclosed during an earnings call in April that the CPU-to-GPU usage ratio for AI inference—the execution of models—was 1 to 4. However, by July, Intel CFO David Zinsner indicated that this ratio had rapidly climbed to nearly 1 to 1. Executives at AMD and Arm have also made similar comments.
Beyond the CPUs themselves, insufficient supply of companion memory chips and physical space constraints within data centers are jointly exacerbating this computing shortage.
Tech Giants’ Battle for Computing Resources
Balancing the internal and external allocation of computing power has become a common challenge for large technology companies over the past two years. Reports indicate that Google established a committee of senior executives last year specifically tasked with coordinating computing resource allocation among Google Cloud, its DeepMind AI research division, and its consumer businesses. Even so, internal friction persists; Noam Shazeer, a star AI researcher at Google, departed earlier this summer partly due to dissatisfaction with access to computing power.
Microsoft’s experience illustrates another dimension: improvements in computing resource management efficiency can directly translate into revenue growth. Microsoft CFO Amy Hood, during an earnings call last month, attributed part of the growth in Azure cloud services to “efficiency improvements” in managing its fleet of CPU and GPU servers. She stated bluntly, “When we are able to drive efficiency gains, those gains get monetized very quickly.”