AI InfrastructureCloud Native

Serverless Inference as a Service: Removing the AI Infrastructure Bottleneck | Jon Alexander, Akamai | TFiR

0

Teams that locked cloud workloads to a single availability zone or region are discovering those early decisions directly constrain how and where AI workloads can run. Unwinding that architectural debt is expensive, slow, and forces engineering teams to focus on infrastructure plumbing instead of business outcomes. The organizations that architected for distributed deployment from the start, even if they launched in a single region, retain the optionality to scale AI across 2, 5, or hundreds of locations without a rewrite.

In this interview on TFiR, Jon Alexander, SVP of Product for the Cloud Technology Group at Akamai, breaks down why architectural decisions made early in the cloud journey determine AI deployment success, and how Akamai’s serverless inference platform removes placement and scaling complexity entirely.

Guest: Jon Alexander, SVP of Product for the Cloud Technology Group at Akamai
Show: TFiR

Here is what every platform engineer and cloud architect needs to know.

Technical Deep Dive

Q: How hard is it to untangle a single-region cloud setup and move to a distributed model?

Jon Alexander, SVP of Product for the Cloud Technology Group at Akamai, frames this as one of the defining decisions of any cloud journey. Organizations that tied their applications to a single availability zone, a single region, or a single database location accumulate compounding architectural debt that becomes extremely difficult to unwind. Teams that planned for multi-location deployment upfront, even if they launched in one region, retain the optionality to scale without rearchitecting.

“If you architect and tie to just one location, you’re gonna have problems, you’re gonna have to unwind a lot of those decisions you made early on.” — Jon Alexander, SVP of Product for the Cloud Technology Group, Akamai

Q: Does the same single-region lock-in problem apply to AI workloads?

Alexander confirms the pattern is identical for AI. The same architectural decisions that created problems for general cloud workloads, tying compute and data to a fixed location, create the same bottlenecks when organizations try to scale AI inference across distributed environments. The advice is consistent: plan for multi-location deployment even if you start in one place.

“I think that the same is true for AI. If you’ve got a plan of how you could scale this out to 2 locations, 5 locations, 10 locations, 2100 locations, you’ll be fine.” — Jon Alexander, SVP of Product for the Cloud Technology Group, Akamai

Q: What is Akamai’s serverless inference platform and what problem does it solve?

Akamai is actively building a serverless inference platform designed so customers can consume AI compute entirely as a service. Teams using this platform do not need to manage where inference is running, how many locations are involved, or handle placement decisions themselves. Alexander identifies serverless as one of the most important infrastructure primitives for AI adoption at scale.

“Customers can just consume as a service and they don’t need to think about where something’s running, how many locations, they don’t need to think about that placement.” — Jon Alexander, SVP of Product for the Cloud Technology Group, Akamai

Q: What is the right architectural approach for teams that are deploying AI in a single region today?

Alexander’s guidance is that teams do not need to build full multi-region infrastructure on day one, but they must have a defined pathway for how they would scale to additional locations. Having that plan preserves optionality later without requiring a full rearchitecture. The absence of that plan is the root cause of the unwind problem teams face as they try to scale AI.

“You don’t need to build all of that infrastructure on day one, but at least having a pathway is going to give you the optionality later on.” — Jon Alexander, SVP of Product for the Cloud Technology Group, Akamai

Resources & Documentation

  • Akamai Cloud Computing, Akamai’s cloud platform including serverless compute and distributed infrastructure offerings

***

👇 Click to Read Full Raw Transcript

Swapnil Bhartiya: How hard is it for organizations to untangle their current setup and move towards a distributed model? And how can Akamai make that easier so enterprises can focus on business outcome and not all the plumbing involved?

Jon Alexander: Again, I think this is one of the kind of key decisions that everyone has as part of their cloud journey as well. So I’m not sure this is different than people have seen before. If you make the right architectural decisions upfront, if you think about where you want to end up, you can architect in the cloud to be able to deploy and run in many locations. But if you make a certain set of decisions early on, if you tie your application to a single availability zone or a single region, or a single database location, then you really end up with challenges. And I think that the same is true for AI. We’ve heavily invested in serverless as one of the kind of fundamental presentations of of compute and infrastructure that we think is going to be incredibly important for AI adoption. We’re working a lot right now in our roadmap on building a serverless inference platform that customers can just consume as a service and they don’t need to think about where something’s running, how many locations, they don’t need to think about that placement. So I think that’s the type of advice I would give is if you architect and tie to just one location,

Swapnil Bhartiya: yeah,

Jon Alexander: you’re gonna have problems, you’re gonna have to unwind a lot of those decisions you made early on. If you’re making decisions where even if you are deploying into one region on day one, but if you’ve got a plan of how you could scale this out to 2 locations, 5 locations, 10 locations, 2100 locations, you’ll be fine. So again, it’s planning ahead. You don’t need to build all of that infrastructure on day one, but at least having a pathway is gon you the optionality later on.

Why AI Infrastructure POCs Stall Before Production and How to Fix It | Rob Hirschfeld, RackN | TFiR

Previous article