Executive summary
Amazon Web Services is experiencing its strongest growth in nearly five years, driven by surging AI demand across enterprise customers and research labs. AWS's AI business has reached a $25 billion revenue run rate, with compute capacity now fully committed through 2028. The company is raising capital expenditure to $220 billion to meet demand that significantly outstrips supply.
What happened
AWS CEO Matt Garman revealed that the cloud computing division is recording its fastest growth rate in nearly five years, powered by widespread adoption of artificial intelligence workloads. The company's AI business has reached a $25 billion revenue run rate, encompassing both model training for major AI labs like OpenAI and Anthropic, and inference workloads for enterprise customers across financial services, healthcare, retail, and media sectors.
Demand has become so intense that AWS has already committed its compute capacity through the end of 2027 and well into 2028, with customers signing five-year contracts to secure access. Amazon is responding by increasing total capital expenditure to $220 billion this year, up $20 billion from prior levels, with plans for continued heavy spending next year.
Garman also highlighted a significant shift in customer usage patterns. Inference, the process of running trained AI models to generate real-world business value, now accounts for roughly two-thirds of all AI compute demand, up from approximately one-third in 2023. This represents a fundamental change in how enterprises are deploying AI, moving from training models to actually putting them to work in production environments.
Why it matters
This development demonstrates that AWS is successfully capitalising on the AI revolution beyond just serving headline-grabbing research labs. The broad-based demand across multiple industries validates Amazon's massive infrastructure investments and provides revenue visibility through long-term customer commitments extending years into the future.
The shift toward inference workloads is particularly significant for AWS's business model. Unlike training, which occurs once per model, inference generates recurring revenue every time a model is used to complete a task or answer a query. Garman's prediction that 80-90% of enterprise AI value will eventually come from inference-powered agents suggests this revenue stream could become increasingly substantial.
AWS's custom silicon strategy, including its Trainium chips for training and Graviton processors, positions the company to capture margin improvements by reducing dependence on third-party chip suppliers. The $25 billion run rate from renting capacity powered by these in-house processors represents a growing competitive advantage in controlling its technology stack.
Bigger picture
The supply-demand imbalance that AWS is experiencing reflects a broader constraint across the entire cloud computing industry. Microsoft and Google have similarly reported capacity challenges, with all three hyperscalers racing to build data centre capacity fast enough to meet AI demand. This tight supply environment strengthens pricing power for cloud providers and validates the enormous capital expenditure commitments being made across the sector.
The shift from training to inference workloads has implications beyond cloud computing. For the cryptocurrency mining sector, which pivoted toward offering AI and high-performance computing services after Bitcoin mining margins tightened following the April 2024 halving, this trend suggests potential challenges ahead. Training workloads, which many mining companies positioned themselves to serve, require different hardware characteristics than inference, which demands lower latency and higher throughput per query rather than massive parallel processing clusters.
The inference dominance also matters for decentralised compute networks like Render, Akash, and io.net, which position themselves as alternatives to centralised cloud providers. While these protocols could benefit from growing inference demand, they face the challenge of matching the low-latency, high-reliability service that inference workloads require, unlike training which tolerates network interruptions.
What to watch
Monitor AWS's quarterly revenue growth rates and operating margin trends to assess whether the strong demand translates into improved profitability as custom silicon scales. Pay attention to whether the company maintains its accelerated capital expenditure pace or begins to moderate spending as new capacity comes online.
Watch for updates on customer contract terms and pricing power indicators. With capacity sold out through 2028, AWS may have leverage to command premium pricing, but competitive pressure from Microsoft Azure and Google Cloud could limit pricing flexibility.
Track the adoption rate of AWS's custom Trainium and Graviton chips versus third-party options like Nvidia GPUs. Higher adoption of proprietary silicon would signal both customer acceptance and potential margin expansion. Also monitor whether the inference-to-training ratio continues shifting toward inference, which would validate Garman's thesis about where enterprise AI value is headed.
Get our top market beating stocks free here
#semiconductors
#ai
#cloud
#capex