Agentic AI systems must make hundreds of autonomous decisions per second across live business processes. Centralized infrastructure built for model training cannot meet that requirement. The latency ceiling that was acceptable for chatbots becomes a hard failure boundary when AI agents replace real-time workflows previously handled by humans or deterministic algorithms.
In this interview on TFiR, Jon Alexander, SVP of Product for the Cloud Technology Group at Akamai, breaks down why the industry’s shift from training-focused deployments to pervasive real-time inference demands a fundamental rethink of AI infrastructure architecture.
Guest: Jon Alexander, SVP of Product for the Cloud Technology Group at Akamai
Show: TFiR
Here is what every platform engineer and AI infrastructure team needs to know.
Technical Deep Dive
Q: Why is inference latency now the defining infrastructure problem for production AI?
Jon Alexander, SVP of Product for the Cloud Technology Group at Akamai, explains that the nature of AI workloads has changed fundamentally. Early deployments focused on training large foundational models in centralized clusters; then came relatively constrained applications like chatbots and customer support. Agentic systems that handle real business processes must operate at machine speed, not human speed, which exposes the latency limits of infrastructure designed for earlier workload types.
“Needing to make decisions at machine speed, not at human speed, is the key transition that we’re seeing.”
Jon Alexander, SVP of Product, Cloud Technology Group, Akamai
Q: How has the progression of AI workload complexity led to the current inference problem?
Alexander traces three distinct phases: large-scale training clusters aggregating data into centralized infrastructure, then simple but capable applications like chatbots and customer support, and now sophisticated use cases including AI coding assistants and, most critically, autonomous agents. Each phase introduced more demanding latency and distribution requirements. The agent phase is where existing centralized infrastructure breaks under real-world conditions.
“We’re on the cusp of a significant change in the types of AI workloads that our customers are looking to deploy.”
Jon Alexander, SVP of Product, Cloud Technology Group, Akamai
Q: What makes agentic AI workloads different from chatbot or coding assistant workloads in infrastructure terms?
Chatbots and coding assistants operate in constrained, relatively one-dimensional interaction patterns. Agents are designed to become pervasive across business processes, replacing algorithms, applications, and human decision-makers across a broad range of tasks simultaneously. That scope requires inference to be continuous, distributed, and real-time rather than request-response at human interaction cadence.
“They’re looking at a much broader range of use cases. This is where inference is evolving and becoming more real time.”
Jon Alexander, SVP of Product, Cloud Technology Group, Akamai
Resources & Documentation
- Akamai Cloud Technology Group, distributed cloud and edge infrastructure for real-time AI inference workloads
***
👇 Click to Read Full Raw Transcript
Swapnil Bhartiya: Let’s talk about latency. Why is it becoming the defining infrastructure problem for production AI, especially as enterprises transition from simple chatbots to complex agentic systems that work autonomously?
Jon Alexander: That’s right, yeah. No, and I think that’s the key point. Like, we’re on the cusp of a significant change in the types of AI workloads that our customers are looking to deploy. And so again, if we think back a couple of years, a lot of the focus was on developing the model. So big training clusters were being developed, aggregating huge amounts of data into centralized infrastructure to create these amazingly powerful foundational models. Then we started to shift into application of AI into relatively simple workloads. Simple in terms of the application, powerful in terms of the capabilities that are being deployed, but primarily things like chatbots, customer support, relatively simple sort of one dimensional interactions that end users were having with AI. Today, one of the most prominent uses for AI that we see is for coding. So we’re starting to get into much more sophisticated use cases, but again, deployed in a relatively constrained environment. What we’re seeing our customers talking about now is moving into more real time AI, where it becomes pervasive throughout their business process. And they’re looking at deploying agents to handle many of the tasks that previously were handled by other algorithms and applications that they had, or even humans. And so they’re looking at a much broader range of use cases. And this is where inference is evolving and becoming more real time. And needing to make decisions at machine speed, not at human speed is the key transition that we’re seeing.





