AI Infrastructure

How to Measure the Unit Economics of AI Inference | Ari Weil, Akamai | TFiR

0

AI workloads introduce a new category of cost that most engineering and finance teams are not yet instrumented to track. Inference requests, session-level data transfer, multi-cloud billing complexity, and LLM licensing all compound into spend that scales faster than the value delivered. Without a clear methodology for measuring cost per use case, organizations cannot determine whether their AI is operating at a sustainable margin or quietly eroding it.

In this interview on TFiR, Ari Weil, VP Product Marketing at Akamai, breaks down how engineering, FinOps, and business operations teams can use real user monitoring, active and passive testing tools, and cost modeling to build measurable, margin-aligned AI architectures.

Guest: Ari Weil, VP Product Marketing at Akamai
Show: TFiR

Here is what every platform engineer and FinOps practitioner needs to know.

Technical Deep Dive

Q: Why can’t most organizations track unit economics for AI inference workloads?

Ari Weil, VP Product Marketing at Akamai, argues that organizations often treat AI as a fundamentally different category of application, which leads them to abandon the cost instrumentation disciplines they already apply to conventional workloads. In reality, an AI capability added to an application still has a measurable size, a cost, and a set of workloads that spawn from it. The gap is not a lack of available tools; it is a failure to apply existing observability, testing, and FinOps methodologies to the AI layer.

“Every business has the tools to do this. It’s just when we look at AI, a lot of times we’re thinking that this special thing somehow completely changes the game.” — Ari Weil, VP Product Marketing, Akamai

Q: How does real user monitoring help model AI inference costs?

Weil describes Akamai CloudTest as an example of a real user monitoring product that captures how many users are hitting an application, where they originate, what devices and network connections they use, and how many requests and responses compose an average session. Those recorded sessions can then be replayed across cloud and edge infrastructure to expose the actual cost dynamics of a given user interaction pattern. This gives teams empirical session-level data rather than estimates as the foundation for cost modeling.

“We can look at the volume of users coming to you, we can look at their recorded sessions and actually show you what that user interaction looks like.” — Ari Weil, VP Product Marketing, Akamai

Q: How do active and passive testing tools decompose AI workload costs?

Weil explains that tools such as Wireshark for passive network inspection, or active testing harnesses that replay recorded user sessions, allow teams to decompose a full user interaction into individual repeatable steps. Each step surfaces the systems touched, the data transferred, and the processing triggered, giving engineers a granular cost map of the workload. Those steps can then be assembled into a testing harness that validates cost and performance assumptions before a workload scales.

“Whether you use a Wireshark to look at the network connections and the network dynamics, or you use another active or passive testing tool to look at those real user activities and replay them and start to decompose them into repeatable steps that you might use in a testing harness.” — Ari Weil, VP Product Marketing, Akamai

Q: What role do FinOps tools play in tracking AI costs across multi-cloud environments?

Weil highlights that multi-cloud environments compound the cost visibility problem because each cloud provider meters and bills resources differently. FinOps tooling normalizes those billing signals so teams can see true workload cost regardless of which provider is running the inference. He recommends combining a FinOps tool with observability platforms and active and passive testing tools as a stack, with Akamai offering several of these capabilities for its customers.

“There are also FinOps tools that you can bring to bear to help you make sense of the cloud bills that you’re seeing, including if you have, or maybe especially if you have a multi-cloud environment where the way that things are being metered and billed are different.” — Ari Weil, VP Product Marketing, Akamai

Q: How should teams model AI costs when they do not yet have full observability instrumentation in place?

Weil notes that even without full tooling, teams can use spreadsheets, databases, and math to model growth curves and extrapolate cost trajectories from what they know about their user interactions and data transfer volumes. The modeling exercise can be virtual or conceptual, and the key inputs are the systems touched, data transferred, and the cost of enabling each use case. This approach gives teams a working cost baseline they can refine as observability matures.

“You can do it virtually or you can do it maybe conceptually in spreadsheets and databases and actually use math. Just to extrapolate out, this is what my growth curves will look like.” — Ari Weil, VP Product Marketing, Akamai

Q: How do you align AI workload cost data with business margin analysis?

Weil frames the end goal as a channel-level margin analysis: once teams know the cost of enabling a specific AI use case, they can work with revenue operations, product, or business operations teams to determine whether the cost is commensurate with the revenue generated and whether the margin profile for that channel or workload is sustainable. This requires breaking AI spend down by use case and channel rather than treating it as an aggregate infrastructure line item.

“By channel you can start to understand with your revenue operations or product or business operations teams, am I incurring a cost that’s commensurate with what I’m making? And does the margin profile for this channel or this workload or this application make sense?” — Ari Weil, VP Product Marketing, Akamai

Q: Does adding an LLM or fine-tuned model fundamentally change application architecture and distribution?

Weil is direct on this point: a frontier LLM or a fine-tuned model added to an application changes the capability being enabled, but it does not fundamentally change the way the application is built and distributed. The AI component is an addition with a size, a cost, and measurable workload activity attached to it. Treating it as a completely different paradigm is what causes teams to abandon the cost discipline they would apply to any other workload.

“It might change the game in the capability that we’re enabling. And it doesn’t necessarily fundamentally change the way that I’m building and distributing an application, it’s just adding a new capability to it.” — Ari Weil, VP Product Marketing, Akamai

Resources & Documentation

  • Akamai, cloud, edge, and security platform with observability, testing, and FinOps capabilities for AI workloads
  • Akamai CloudTest, real user monitoring and session replay product for modeling application and AI workload costs
  • Wireshark, open-source passive network protocol analyzer for inspecting network-level workload dynamics

***

👇 Click to Read Full Raw Transcript

Swapnil Bhartiya: Now let’s talk about tracking and measuring. If most organizations can’t even track unit label economics for inference, how do they even know whether their AI is scaling efficiently or just getting more expensive?

Ari Weil: Yeah, so I think in some cases I’ll use an example from Akamai. We have a product called CloudTest. That product is designed to look at the real user monitoring of all of the users of your application, your SaaS application in the case of a proxied app on Akamai, and to show you how many people are using it, where are they coming from, what type of devices and network connections are they, how many requests and responses are part of an average user session. And then you can start to replay those things through your cloud and your edge network so that you understand exactly the dynamics. Whether you use a Wireshark to look at the network connections and the network dynamics, or you use another active or passive testing tool to look at those real user activities and replay them and start to decompose them into repeatable steps that you might use in a testing harness. But there are multiple ways to get whether you’ve architected the whole user interaction or you’re just recording the reality of a user interaction coming to you, all of those individual steps, the systems that you hit, the amount of data that you typically transfer, and then it becomes a modeling exercise. You can do it virtually or you can do it maybe conceptually in spreadsheets and databases and actually use math. Just to extrapolate out, this is what my growth curves will look like. You can use real user examples like that CloudTest product that I mentioned, so that we can look at the volume of users coming to you, we can look at their recorded sessions and actually show you what that user interaction looks like. And there’s a myriad of other observability tools that can help you do something similar. There are also finops tools that you can bring to bear to help you make sense of the cloud bills that you’re seeing, including if you have, or maybe especially if you have a multi cloud environment where the way that things are being metered and billed are different. So I would take my FinOps tool, I would take my observability tool, I would take my active and passive testing tools, and if you’re an Akamai customer, we have a number of these for you. You can also use services if you want to, to help you with sort of this modeling and observing your actual user workloads. But then it basically becomes models and math. It’s how much does it cost for me to enable this use case. Is that use case as optimized as I might want it to be? How much does that cost me? And then by channel you can start to understand with your revenue operations or product or business operations teams, am I incurring a cost that’s commensurate with what I’m making? And does the margin profile for this channel or this workload or this application make sense? And I think every business has the tools to do this. It’s just when we look at AI, a lot of times we’re thinking that this special thing that is a frontier LLM that maybe we’ve licensed, or maybe it’s a fine tuned or a post trained model that we’ve created ourselves, somehow completely changes the game. It might change the game in the capability that we’re enabling. And it doesn’t necessarily fundamentally change the way that I’m building and distributing an application, it’s just adding a new capability to it. But that capability has a size and it has a cost and it has a set of activities that workloads spawn from it or to it. And you can measure and meter all of those things to come up with the right architecture and the right scalability model for your business.

What Changes in Infrastructure When Every App Uses AI | Danielle Cook, Akamai | TFiR

Previous article

How to Assess Failover Health in 5 Minutes Using Logs | Alexus Gore, SIOS Technology | TFiR

Next article