Token costs do not scale linearly. Routing every task through a multi-trillion-parameter model feels safe in development and becomes financially catastrophic in production. Most teams discover this too late, after architecture decisions are already locked in.
In this interview on TFiR, Dr. Robert Blumofe, Executive Vice President and Chief Technology Officer at Akamai, breaks down why the brute-force approach to AI infrastructure fails at scale and what a disciplined, cost-aware architecture actually looks like.
Guest: Dr. Robert Blumofe, Executive Vice President and Chief Technology Officer at Akamai
Show: TFiR
Here is what every platform engineer and AI infrastructure architect needs to know.
Technical Deep Dive
Q: Is the current AI infrastructure model sustainable at scale?
Dr. Robert Blumofe, Executive Vice President and Chief Technology Officer at Akamai, argues the brute-force approach to AI infrastructure is not sustainable as applications move toward production scale. The core problem is that teams default to large, general-purpose models for every task regardless of whether that task requires that level of capability. The compounding token costs and infrastructure overhead make real-world scalability extremely difficult without a more deliberate architectural strategy.
“I get the temptation to use this brute force approach. On the small scale maybe it’s okay. But if you tried to scale that to a real application that’s going to be used by millions of people, using the wrong model is just a killer.” — Dr. Robert Blumofe, Executive Vice President and Chief Technology Officer, Akamai
Q: What is the alternative to the brute-force approach to AI infrastructure?
Blumofe describes the alternative as an intelligent approach built on three principles: using the right AI model for the task, using non-AI tooling whenever possible because it is cheaper, and delivering infrastructure in the right place. Applying all three together dramatically lowers cost and makes AI applications scalable. He frames this not as a radical shift but as disciplined engineering applied to decisions teams are already making.
“Use the right AI for the task. You don’t have to do everything with an ask-me-anything multi-trillion-parameter model. Use non-AI whenever you can, because that’s much cheaper.” — Dr. Robert Blumofe, Executive Vice President and Chief Technology Officer, Akamai
Q: When should teams use non-AI tooling instead of an LLM?
Blumofe’s position is direct: use non-AI tooling whenever it can accomplish the task, because non-AI is significantly cheaper than invoking a language model. The decision should be driven by task requirements, not by a default preference for AI. Over-indexing on LLMs for tasks that do not require reasoning or generation inflates token costs without delivering proportional value.
“Use non-AI whenever you can use non-AI, because that’s much cheaper.” — Dr. Robert Blumofe, Executive Vice President and Chief Technology Officer, Akamai
Q: How does model selection directly affect AI application scalability?
Blumofe uses his own experience with tools like Open Clon and Hermes Agent to illustrate the problem. He found himself configured to use Claude Opus 4.7 for tasks that did not require a model of that capability, generating token fees that were disproportionate to the work being done. At individual scale that is a minor inconvenience. At application scale, with millions of users, the wrong model selection becomes a direct threat to the viability of the product.
“I’m racking up these ridiculous token fees. Okay, it’s one thing for me to spend a little bit more money personally for my own use. But if you tried to scale that to a real application used by millions of people, using the wrong model is just a killer.” — Dr. Robert Blumofe, Executive Vice President and Chief Technology Officer, Akamai
Q: Does AI infrastructure need AGI or quantum computing to become sustainable and ubiquitous?
Blumofe is explicit that neither AGI nor quantum computing is required for AI to become ubiquitous and sustainable. AI as it exists today is sufficient to deliver significant improvements to the experiences people have using computers and digital services. What is required is good engineering discipline and intelligent infrastructure choices applied to the models and tools already available.
“We don’t need any fancy new breakthroughs, we don’t need AGI, we don’t need quantum computing. AI as it lives today, with some good engineering and some good intelligent choices, can deliver some really phenomenal up-levelings of the experience that we all have.” — Dr. Robert Blumofe, Executive Vice President and Chief Technology Officer, Akamai
Resources & Documentation
- Akamai, cloud and edge infrastructure platform referenced throughout the discussion on delivering AI in the right place
***
👇 Click to Read Full Raw Transcript
Swapnil Bhartiya: The current state of AI doesn’t seem very sustainable. From the massive energy demands to token costs going through the roof, looking two to three years ahead. How does AI infrastructure need to evolve to actually become sustainable?
Dr. Robert Blumofe: I really do think it comes down to sort of an intelligent sort of alternative approach to the brute force approach. And the intelligent approach is actually fairly simple and it really is the things that we were just talking about, it’s what you build your agent. It’s use the right AI for the task. You don’t have to do everything with an ask me anything multi trillion parameter model. It’s using the right tools for the task, right use non AI whenever you can use non AI because that’s much cheaper. It’s delivering the right infrastructure to the task and delivering it in the right place. Doing all those things together can dramatically lower the cost and make these applications scalable and therefore much more usable. And I get the, I get the temptation to sort of, you know, use this brute force approach. And on the small scale maybe it’s okay. You know, anecdotally, you know, I’ve been, I like to play with these agents and play with LLMs and I’ve been using things like Open Clon and Hermes Agent and I oftentimes find myself, I’ve got the thing configured to use Claude Opus 4.7 for example, which is a great, just a phenomenally great model. But then I’m sort of looking at the stuff that I’m doing with it thinking wait a minute, you know, I don’t need that level of model to do what I’m doing. So I’m racking up these ridiculous, you know, token fees and you know, okay, it’s one thing for me to, you know, spend a little bit more money, you know, personally just for my own use, but if you tried to scale that to a real application that’s going to be used by, by millions of people, using the wrong model is just a killer and using the wrong infrastructure is just a killer. So you’ve got to have the right models, the right tools, the right infrastructure in the right place. That intelligent approach is what makes the whole thing scale and is what ultimately is going to make AI ubiquitous. And, and it is going to be ubiquitous, you know, and we don’t need any fancy new breakthroughs, we don’t need AGI, we don’t need quantum computing. AI as it lives today, with some good engineering and some good intelligent choices can deliver some really phenomenal up levelings of the experience that we all have using computers or using any services.





