Why Neo-Clouds and the CPU Are Quietly Winning the AI Inference Race

August 19th, 2026 | | 4:15
Share this post:
Facebook | Twitter | Google+ | LinkedIn | Pinterest | Reddit | Email
 
This post can be linked to directly with the following short URL:


 
The video player code can be adjusted to different sizes:


 
This video file can be linked to by copying the following URL:


 
Right/Ctrl-click to download the video file.
 
Subscribe:
Connected Social Media - iTunes | Spotify | YouTube | Twitter | RSS Feed
Tech Barometer - iTunes | Spotify | RSS Feed
 

As AI shifts from training to inference and call-and-response chatbots give way to autonomous agents, AMD’s Brayden Mahdavi explains why enterprises are rebalancing the CPU-to-GPU ratio and turning to nimble neo-clouds for cheaper, faster compute.

Get tech leader insights to move faster and smarter.

Get more stories by subscribing to The Forecast.

Video transcript:

Brayden Mahdavi: I really sit at the intersection of three different groups. I sit between cloud providers, so this is your kind of typical cloud service provider, then your infrastructure providers, and then end customers within our enterprise. My role really gives me a lot of energy because I hear a common theme across all of these groups, and that is customers want options. I’m really seeing it across three main layers. So the first is really the infrastructure layer, right? This is your silicon choice and optionality across different silicon providers. And then you have within the cloud service provider layer, right? So my traditional tier one hyperscalers, but now we’re seeing this new emergent AI specialized cloud provider often referred to as neo-clouds. And the third layer is really your virtualization layer. As many enterprises know, licensing structures have changed as of recent, and that is creating a lot of challenges and a lot of ingenuity around how we reinvent our stack to better serve our business needs.

If many enterprises can put their whole fleet on this tier one cloud, albeit Amazon, Google, Microsoft, and that business is critical for AMD. It’s not going away anytime soon. Given some of the challenges that we’re facing in 2026, capacity and supply are huge challenges. We are really hungry at a market level for a new offering, right? And this is where you are seeing these emergent AI cloud providers that are specialized historically maybe in the GPU space, but many of our partners are seeing with the shift to agentic AI, a bigger need for a CPU. They’re so nimble and lean that they can help productize our technology in a way that is favorable for our end customers who are facing challenges around total cost of ownership and other performance challenges. We are seeing a major industry shift from the classic call and response AI that was your traditional ChatGPT to now agent calls.

And agents can call more and more agents to do more autonomous tasks for us. They bring us a lot of value. What we don’t see is that they demand a lot of compute resources. What we are going to see very likely over the next few years is a shift back to parody, if you will, of the CPU to GPU ratio. So where our enterprise customers were typically buying as many as eight GPUs to one CPU, we’re seeing that ratio come down much more to a one-to-one ratio because our CPUs are able to handle much of the orchestration and much of the agent calls, the tool calls that are required for these types of workloads. The last few years from that ChatGPT moment in 2023, when a lot of people realized this is a real invention that is upon us today, the shift has really been many training cycles within the frontier lab environments.

We’re doing a lot of forward passes and back propagation at scale to train up these models, whereas in the next couple of years, it will be very heavily favored towards inference. Now we want to see the outputs of these models. And as we get to that shift, the demand for compute just continues to soar through the roof. Our customers are starting with their workload need and working back to solve what the right infrastructure stack is. We are seeing more optimization in that process, whether it’s through a FinOps practice or whether it’s through just a straight up business need. It’s at a time where we are moving into optimization and how do I lower my total cost of ownership? And between AMD and our partnerships with Nutanix and with our AI cloud providers, we are very confident that we are going to be able to help our end customers accomplish their business requirements.

Transcript Read/Download the transcript.
 

Tags: , , , , , , , , , , , , ,
 
Posted in: Artificial Intelligence, Cloud Computing, Tech Barometer - From The Forecast by Nutanix, Video Podcast