The boundary between AI applications and traditional applications is gone. Every legacy system, SaaS platform, and internal workflow is being absorbed into AI pipelines through APIs, RAG connections, and agent chains, creating infrastructure demands that most architectures were never designed to handle. The result is compounding latency, surprise egress bills, and production systems that behave unpredictably under load.
In this interview on TFiR, Danielle Cook, Director of Product Marketing at Akamai, breaks down why every application now participates in the AI ecosystem, how to architect for data in motion from day one, and what platform and infrastructure teams must get right before production AI exposes the gaps.
Guest: Danielle Cook, Director of Product Marketing at Akamai
Show: TFiR
Here is what every platform engineer, infrastructure architect, and AI application team needs to know.
Technical Deep Dive
Q: Is the distinction between AI applications and non-AI applications still meaningful for enterprises?
Danielle Cook, Director of Product Marketing at Akamai, says the line has disappeared entirely. All applications are now part of the AI ecosystem, either by connecting to AI through APIs, feeding data into models for training, or supporting workflows that do. A legacy ERP system that has not changed a line of code becomes an AI application the moment someone points a RAG pipeline at it.
“You can’t really opt out of the AI ecosystem anymore. You’re in it.” — Danielle Cook, Director of Product Marketing, Akamai
Q: How much visibility do organizations actually have into whether their code is touching AI?
Cook says visibility depends on organizational maturity. Regulated industries can lock things down, but any organization with external code contributions or active developer workflows should assume AI is already in use, approved or not. Shadow AI is becoming a systemic challenge, and platform teams are struggling to monitor the volume of AI-generated code flowing into production.
“Shadow AI is going to keep growing. People are just using it, whether it’s approved or not.” — Danielle Cook, Director of Product Marketing, Akamai
Q: What changes when an enterprise moves AI from experimentation to production?
In experimentation, the focus is the model: does it work, is the output good. In production, real users, real latency requirements, real data, and real bills enter the picture. Because inference is continuous rather than a one-time training run, GPUs are only part of the architecture. Compute, storage, and networking all determine how an application performs, and throwing more GPUs at every bottleneck is not economically viable.
“Choosing the right model might be an exciting decision, but making everything work around it at production scale is an infrastructure problem.” — Danielle Cook, Director of Product Marketing, Akamai
Q: Does AI require a complete rethink of infrastructure architecture or just incremental adjustments?
Cook says architects need to evaluate the entire infrastructure stack. AI workloads are real-time, interactive, and involve constant data in motion, which changes the cost and design calculus at every layer. Architects must account for model routing, semantic caching, compute placement relative to users, and the economics of data movement, not just GPU provisioning.
“How do you build for that while keeping it cost effective? How do you not blow your entire budget trying to deliver an amazing experience?” — Danielle Cook, Director of Product Marketing, Akamai
Q: Why is moving compute to data, rather than data to compute, becoming the dominant paradigm for AI workloads?
Cook explains that in AI deployments the model, the data, and the user can exist in three separate locations. Data is often the least movable element, constrained by size, sensitivity, regulation, or the cost of transfer. Workload placement decisions now hinge on latency requirements, data location, compliance jurisdiction, and egress economics, and those decisions belong in the architectural conversation from the start.
“If your cloud economics assume data is mostly staying put but your AI architecture assumes it’s constantly moving, those assumptions are going to collide.” — Danielle Cook, Director of Product Marketing, Akamai
Q: What is actually happening across infrastructure when a user sends one request to an AI-enabled application?
What looks like a single request to the user can trigger dozens of distributed requests underneath, hitting multiple models, vector databases, object storage, APIs, agents, and GPU infrastructure that are not necessarily co-located. A small amount of latency repeated across 50 hops is no longer a small amount of latency. The bottleneck may be the network, an API, a storage layer, or a single step in a long agent chain, not the model or GPU.
“The old application architecture you fit on a napkin. Today’s architecture is messy, and it gets messy really quickly.” — Danielle Cook, Director of Product Marketing, Akamai
Q: How should data residency, regulatory jurisdiction, and multi-cloud distribution influence where AI compute runs?
Cook says the starting point is understanding where data lives and where it is permitted to live. Geography affects performance, compliance, resilience, and cost simultaneously. Data transfer charges, including egress fees Cook notes are estimated at 70 to 80 billion dollars annually across the industry, can fundamentally change the economics of an AI architecture if not accounted for during design rather than after deployment.
“You don’t want to be discovering your egress bill after you’ve designed your architecture. You want to be designing for it up front.” — Danielle Cook, Director of Product Marketing, Akamai
Q: If you were designing an AI application platform from scratch today, what would you build differently to avoid future refactoring?
Cook would design for data in motion from day one, assuming the application, model, data, and compute will span multiple clouds and regions rather than a single environment. She would make workload placement flexible so compute can be positioned based on end-user experience requirements, and prioritize model and infrastructure portability given how rapidly the model landscape is evolving. The goal is to absorb new model releases, new regions, and new compliance requirements without redesigning the application.
“If the application is distributed by design, the infrastructure underneath it should be too.” — Danielle Cook, Director of Product Marketing, Akamai
Q: Where should teams re-architecting for AI start, beyond data placement?
Cook outlines several immediate considerations: available compute resources and who controls them, which models to run and at what size, how much of the data and infrastructure the organization wants to own versus consume as a service, and the tokenomics of each approach. Teams must also decide whether they want to engage with infrastructure at all, since vendors are increasingly offering paths to AI inference that abstract the infrastructure layer entirely.
“AI is expensive. How much do you want to spend on it yourself versus paying others? That has to be part of the conversation.” — Danielle Cook, Director of Product Marketing, Akamai
Q: Do larger models always deliver better results, and how should teams think about model selection for cost and performance?
Cook is direct: bigger is not always better. Running a fine-tuned small model on a CPU for a specific customer interaction can deliver an equivalent or better experience at significantly lower cost than routing the same request to the largest available model. Teams need to understand their workloads well enough to route each request to the most cost-effective compute, whether CPU or GPU, and resist the instinct to default to maximum model size.
“You need to understand your workloads, route to the most cost-effective CPU or GPU, and stop wasting money on the assumption that bigger is always better.” — Danielle Cook, Director of Product Marketing, Akamai
Q: Where does edge computing fit into the AI infrastructure picture?
Cook says edge compute at Akamai is about ensuring customers can service their end users with the right compute in the right location. That can mean distributed GPU locations, containers at the edge, or other configurations depending on the experience required and the financial model that supports it. The options vary based on what the customer needs to deliver and what cost structure is acceptable.
“For us, it’s about making sure that compute happens where customers need it.” — Danielle Cook, Director of Product Marketing, Akamai
Q: How should organizations evaluate their AI strategy to ensure they are solving real problems rather than chasing hype?
Cook reframes the question around a framework rather than a maturity model. Some organizations are well served by consuming an AI service and layering security and bot management on top. Others need their own GPUs, Kubernetes, fine-tuned models, and custom storage configurations. Some want to train on their own data at scale. Others want serverless inference with no infrastructure ownership at all. The right position is determined by business objectives, revenue targets, and financial model, not by a desire to reach a higher maturity tier.
“People have to remind themselves to look at what the business objective is. Does this map to what you’re trying to achieve? It isn’t one size fits all.” — Danielle Cook, Director of Product Marketing, Akamai
Q: How is democratized AI infrastructure through platforms like Akamai Cloud changing how developers work?
Cook says Akamai Cloud’s mission is to make the developer experience simple enough that teams can access Kubernetes clusters, Linux clusters, and GPU clusters without infrastructure expertise becoming a prerequisite. The same simplification is being applied to agentic workloads and AI inference, with the goal of letting developers focus on building applications rather than managing infrastructure plumbing.
“When you sign in to Akamai Cloud, you get what you need and you don’t need a degree in understanding how to deploy everything.” — Danielle Cook, Director of Product Marketing, Akamai
Resources and Documentation
- Akamai Cloud, distributed cloud platform offering GPU clusters, Kubernetes, Linux clusters, and AI inference infrastructure
***
👇 Click to Read Full Raw Transcript
Swapnil Bhartiya: For most enterprise IT, AI used to live in its own lane. There was a dedicated model, dedicated app, and dedicated team. That lane is now gone. AI is slipping into every application, every workflow, and every layer of infrastructure. And the moment that happens, decisions about performance, cost and data movement stop being niche choices and they become architectural decisions. This shift is felt more by the companies that never thought of themselves as AI companies at all. Danielle Cook, Director of Product Marketing at Akamai, has been watching this transformation from the inside. She joins me to explore why the boundary between AI applications and everything else has collapsed and what it means for how we build and run applications. Next. Danielle, it’s great to have you on the show.
Danielle Cook: It’s always great to be with you.
Swapnil Bhartiya: A few years ago, you could draw a clean line between AI applications and everything else. Today, every line of code goes through AI in some form. So is that distinction still holds true? Is it still meaningful? And what about those companies that don’t think of themselves as AI companies? What does it mean technically for those companies?
Danielle Cook: So I don’t think you can really opt out of the AI ecosystem anymore. You’re in it. So where we used to be able to say there are two buckets of applications, there’s an AI bucket and a non-AI bucket, that line has disappeared. So all applications contain AI. They’re either connecting to AI through APIs, they’re feeding data into AI to help train models, or they’re supporting something that does. So even if you’re thinking of a legacy ERP system, for example, that’s becoming part of an AI workflow somewhere, as soon as somebody puts a RAG pipeline pointed at it. So nothing about that application changed, but it is in the AI ecosystem. So I think we need to start thinking about, instead of asking which of my applications is an AI application, we need to start focusing on how are my applications participating in the whole AI ecosystem. Because you don’t have to build AI to become part of this system, this ecosystem.
Swapnil Bhartiya: How much visibility do organizations actually have into whether their code is touching AI at all?
Danielle Cook: It depends on the maturity of the organization. Of course, to your point, regulated industries, they’re able to kind of lock things down, but if you are having any code contributed, any workflows, you have to assume people are using AI. I think we’re going to hear shadow AI more and more and more. People are just using it, whether it’s approved or not. So it’s going to be increasingly important to get that visibility and understand where your entire team is. And I know platform teams are struggling with how do we monitor this. Developers are generating so much code, and how do you check it. It’s a real struggle across the industry.
Swapnil Bhartiya: When enterprises move from experimenting with AI to running it in live applications, what actually changes? Why does production AI put different demands on infrastructure as compared to running traditional applications?
Danielle Cook: In production you tend to focus on, or in experimentation rather, you tend to focus on the model. Like, does it work? Is my output good? But in production now, you have real users, real latency requirements, real data, real bills that are associated with this. And because inference is continuous, it’s not a model that you train and walk away from. That’s where you start realizing that GPUs are only part of the architecture. You need compute, you need storage, you need networking, all of that to actually determine how your application performs. And you can’t just throw more GPUs at this at every step. So your infrastructure is changing. You also might be looking at it going, okay, so GPUs, I need it for this step, but actually I need tool calls, APIs, databases, all of that can actually largely work with my CPUs. So choosing the right model might be part of an exciting decision for you, but making everything work around it at production scale is an infrastructure problem that organizations are having to solve.
Swapnil Bhartiya: How much is this changing how architects approach infrastructure? Is it a minor adjustment or do they need to completely rethink their infrastructure strategy?
Danielle Cook: Architects should be looking at the entire infrastructure. AI does change everything. It is real time, it’s interactive, it is the data moving all around, complete data in motion. So how do you build for that while maintaining it, making sure it’s cost effective? How do you not blow your entire budget trying to deliver an amazing experience to your customers? And so when you’re architecting for it, you have to be considering things like model routing. You have to be considering how you’re going to use semantic caching so that the right prompts, the right answers, the right experience gets to your customers. And so yes, it’s an infrastructure challenge. You need to be putting compute close to your users, but you also need to be thinking about the core data movement around inference, around AI, about the customer experience.
Swapnil Bhartiya: With token cost latency and performance pressures rising, moving AI workloads to the edge is becoming critical, especially since not everyone can run things locally. So what paradigm shifts are you seeing in how technology, hardware and organizations like Akamai are positioning AI workloads at the edge to cut down on cost and performance impact.
Danielle Cook: So we used to move data to the compute, right? And increasingly AI is forcing us to move the compute to the data. So we’re seeing the model, the data, the user, those can all be in three separate places. And data is often the most, least movable part of that equation. So whether it’s because of the size of it, sensitivity, regulation, or simply it could be just the cost of moving that data. So we’re seeing that workload placement is becoming a huge decision and it’s a decision around latency, data location, regulations, and the economics, to your point, around tokenomics. And so this absolutely belongs in the architectural conversation, because AI applications, that data again is constantly moving. So if your cloud economics assumes that data is mostly staying put, but your AI architecture assumes it’s constantly moving, those assumptions are going to collide. So you don’t want to be in a situation where you’re discovering your egress bill after you’ve designed your architecture. You want to make sure that you’re designing for it up front, especially with these real time customer interactions.
Swapnil Bhartiya: When someone sends a simple request to an AI-enabled app, what is actually happening across the infrastructure? Where do these systems actually run? And this is a challenge for both teams running those applications and for companies like Akamai who are building that infrastructure for those teams. So can you talk about what’s the reality behind one simple request?
Danielle Cook: What looks like one request to an end user, whether that’s internal or an external customer, it can be dozens of distributed requests underneath. So a user sees a prompt and a response. Underneath that request they might hit multiple models, a vector database, object storage, APIs, agents, GPU infrastructure. Like all of those things aren’t necessarily sitting next to each other. And so that’s why this networking matters so much. A little bit of latency repeated 50 times, that’s not a little bit of latency anymore. And so it means that the bottleneck isn’t necessarily the model or the GPU. It could be the network, it could be an API, it could be storage, it could be one step in a long agent workload. It’s again a very different architecture from the traditional app to database and back model. So I think the big takeaway here is the old application architecture you fit on a napkin. You could draw request response. But today’s architecture, it’s messy. And it gets messy really quickly. Managing that, having to be able to see what’s going on, routing it, understanding it, observing it, are all challenges that developers, platform teams, they’re all experiencing today.
Swapnil Bhartiya: Enterprise data lives across public cloud, private infrastructure, SaaS platforms and edge locations with regulatory requirements spanning different jurisdictions. How should that reality influence where AI compute runs and how much weight vectors like token costs should have?
Danielle Cook: I think that really starts where you need to consider where the data lives and where it’s permitted to live. Because moving that, it can be impractical, slow, there could be multiple restrictions that get in the way. And so geography will impact performance, compliance, resilience and cost. We know that data transfer charges can change the economics. Egress fees are estimated to be around 70 to 80 billion dollars annually. It’s expensive. So when you are moving data, considering it across different clouds, you need to consider how you’re going to manage if you are leaving providers, what is the cost economics of doing that and making sure that it’s sensible for your business. And that’s why you need to be architecting for those egress fees up front and understanding all of those performance restrictions, requirements, all of that, so that you’re safeguarding against that. Because AI is only going to grow.
Swapnil Bhartiya: If you were designing an application platform today for where AI is headed, not where it is, what would you build differently from the start to avoid refactoring, migration or the challenges organizations are facing today?
Danielle Cook: I would absolutely design for data in motion from day one. So I wouldn’t assume that the application, model, data, compute, any of that is going to live in one cloud or one region. I’d assume applications are distributed and that there’s constant data in motion. I’d make workload placement flexible so that I can decide where the compute needs to service my end user based on what their experience should be. And portability fundamentals. So models, infrastructure, all of this is changing rapidly. We see new models coming out continuously. So I’d want to make sure that I can move to what’s right quickly. And my goal wouldn’t be to introduce a new model, region, or compliance requirement without redesigning the application, or to be able to introduce all of that without redesigning the application. So I think that’s really where a distributed infrastructure becomes interesting, because if the application is distributed by design, the infrastructure underneath it should be too.
Swapnil Bhartiya: Now let’s look at teams who are re-architecting now. Where should they start? You mentioned data, what else should be on their priority list.
Danielle Cook: Well, I think you need to obviously look at the compute resources you need. Who has those? What’s available to you? You need to be considering your models, what models you want to run and what size, what amount of the data do you want to own? Do you want to be running your own infrastructure, owning your own model, or do you want to be using another person’s tools, like the build versus buy argument and the tokenomics of it all. So we know that AI is expensive and so it’s the amount of money you want to spend. Do you want to be spending it for yourself? Do you want to be paying it to others? I think the other thing that, as we kind of move into this world, it’s whether you want to even care about the infrastructure. Do you want to consider that there might be ways that you can service your application, your AI application, without ever having to consider the infrastructure? I think we’re going to see, we see vendors doing that, we’re going to see more vendors coming to market with that.
Swapnil Bhartiya: We often get obsessed with the biggest model possible. But smaller models sometimes deliver better results. Can you clarify that?
Danielle Cook: I want to use the biggest model, the most powerful model, but that’s not always cost effective and it’s also not always the best experience for your customer. So being able to understand that I have trained a model and now I’m fine tuning or now I’m servicing this customer, but I actually only need to run this on a CPU with a small size model and they’re going to get a great experience and I’m going to save money or I’m going to do this more cost effectively is really powerful. You need to understand your workloads, you need to be able to route to the most cost effective CPU or GPU, and you need to not waste money on just going bigger is always better.
Swapnil Bhartiya: Where does edge computing fit into the AI infrastructure picture you are describing?
Danielle Cook: So for Akamai, edge compute is obviously something that is near and dear to us. We want to make sure that our customers are able to service their customers with the compute they need. And so that might look like our distributed GPU locations, it might look like using containers at the edge. There’s so many different options that a customer can have based on what their customer needs to do for their customer, what the experience they want, and also the financial impact on that. So for us, we’re about making sure that compute happens where customers need it. And so we have multiple ways of doing that.
Swapnil Bhartiya: As much as I love AI, I also feel that we are in an AI hype cycle, especially around AGI. But honestly, it’s more about GPU count than actual AGI. Every company labels almost everything they do as AI. So when an organization is looking at production AI, what should they actually focus on? Of course infrastructure has been democratized. So real question is, are you solving a real problem using AI or are you just chasing the next shiny object, which is AI. What is your advice to companies? How should they look at their own AI strategies?
Danielle Cook: I think that’s fair. The way I like to think about it is in terms of there’s an AI framework that exists, not a maturity model. Not everyone’s at different stages and so it might be right for a business to just use an AI service and they’re good to go and they’re happy to spend the money on that. And all that they’re putting on top of that is security, managing bot traffic, all of the things that they need to do, but they’re quite happy with that. It might be that they want their own GPUs, they want to do some fine tuning of a model. Now they need to be considering, well, what GPUs do I need? Am I going to use Kubernetes? What is the storage I’m going to put behind this? And they have to be putting that all together. For some organizations it might be, we want to train our own data, we need thousands of GPUs and that’s what’s right for them. And then you have the other end of the spectrum. I don’t really care about infrastructure at all. I don’t want to worry about it. I don’t want my developers to worry about it. I just want to be able to access inference. And then you have serverless inference coming at the end. So it isn’t necessarily I’m going to go here and then graduate to this stage and mature to that stage because I think it’s what’s right for the business. And people have to remind themselves to look at what is your business objective. Like, does this map to what you’re trying to achieve? Is it the right financial model, revenue targets, all of that? Are you spending too much on tokens or do you need to bring it in house? All of those things have to be considered because it isn’t one size fits all.
Swapnil Bhartiya: As AI is being democratized through infrastructure like Akamai Cloud where you can get Kubernetes clusters, Linux clusters, GPU clusters in seconds without having to worry about managing them yourselves, how is this also allowing developers to focus on building business applications instead of all the plumbing that wastes most of their time?
Danielle Cook: It’s absolutely our mission at Akamai Cloud to make sure we are a simple experience for developers. So when you sign in to Akamai Cloud you get what you need there and you don’t need a degree in understanding how to deploy everything. As we support agentic workloads, AI inference, we’re doing the same thing. We’re making it really easy to come into Akamai Cloud and see the options and just get started. So we’ll have plenty of news coming out on that soon so you’ll be able to see it for yourself.
Swapnil Bhartiya: Danielle, thank you so much for taking time out and such a clear-eyed look at where AI is taking application architecture, how organizations should prepare themselves for AI. Thank you for all those great insights and those who are watching please head to akamai.com and learn more about the work they are doing in the AI space. Thank you so much. It was a great discussion and I look forward to having you back on the show for another great conversation. Thank you for your time today.
Danielle Cook: Thank you so much.





