AI Infrastructure

Why AI Agents Waste Tokens on Every Query | Aaron Kao, Pinecone | TFiR

0

Every time an enterprise AI agent answers a question, it re-assembles context from scratch: scanning unstructured documents, ticket systems, CRM logs, and data warehouses in real time. Roughly 85% of LLM token spend goes toward that search step alone, before any useful reasoning begins. The result is high costs, slow responses, and accuracy rates that stall at 50 to 60% task completion because the agent is guessing from incomplete, expensively assembled context.

In this interview on TFiR, Aaron Kao, VP of Marketing at Pinecone, breaks down how Pinecone Nexus, a bring-your-own-cloud knowledge engine for enterprise AI agents, eliminates per-query context compilation by building a governed knowledge layer once and serving it across every agent call.

Guest: Aaron Kao, VP of Marketing at Pinecone
Show: TFiR

Here is what every platform engineer and AI infrastructure team needs to know.

Technical Deep Dive

Q: Why do enterprise AI agents burn so many tokens before returning an answer?

Aaron Kao, VP of Marketing at Pinecone, explains that current agentic architectures rebuild context on every query by searching across databases, unstructured documents, ticket systems, and SaaS platforms simultaneously. Approximately 85% of LLM token consumption is spent on that context search phase rather than on the actual reasoning task. Because every agent call repeats the same search across the same corpus, costs compound rapidly across users and use cases.

“Around 85% of your LLM tokens is spent towards just searching context across lots and lots of data.”

Aaron Kao, VP of Marketing, Pinecone

Q: What is Pinecone Nexus and how does it change the agent retrieval model?

Nexus is a knowledge engine that compiles enterprise data, including unstructured documents, structured databases, and workflow metadata, into what Pinecone calls knowledge artifacts: entity relationships, summaries, and resource maps built ahead of query time. Instead of re-compiling context on every agent call, the knowledge layer is built once and queried repeatedly. Kao frames the model as “compile once, serve many times,” which is the inversion of how agentic RAG currently works.

“Instead of that context being compiled every time at query time, we do it up front at compile time. And then your agent is able to query that very effectively, very efficiently and move very fast.”

Aaron Kao, VP of Marketing, Pinecone

Q: What measurable performance gains does Nexus deliver over agentic RAG?

Based on Pinecone’s benchmarks and early access customers, Nexus delivers 90% fewer tokens consumed compared to standard agentic retrieval, answers up to 30 times faster, and task completion rates rising from 50 to 60% to over 90%. Kao attributes these gains to eliminating the per-query search step across large, heterogeneous data sources including Google Drive, Asana, Gong, and Salesforce.

“We’re actually on our benchmarks able to see 90% fewer tokens, up to 30 times faster answers and task completion from like 50 to 60% to close to over 90%.”

Aaron Kao, VP of Marketing, Pinecone

Q: What is KnowQL and why do agents need a structured query language?

KnowQL is Pinecone’s structured query vocabulary for agents, analogous to SQL for application databases. Current agent data access is unstructured: agents issue natural language queries and receive unfiltered responses with no mechanism to specify filters, provenance requirements, confidence thresholds, output shapes, or token budgets. Kao says KnowQL introduces structured queries and structured responses to the agent layer, making retrieval deterministic and auditable.

“We built KnowQL as a way to have structured queries and structured responses back. That’s not possible in the current agentic world.”

Aaron Kao, VP of Marketing, Pinecone

Q: How does Nexus protect enterprise proprietary data and trade secrets from leaving the organization?

Nexus is deployed as BYOC, bring your own cloud, meaning the entire data plane runs inside the customer’s own VPC or cloud environment. Pinecone does not touch, see, or access customer data. Kao notes this model was a direct response to enterprise concern about frontier model providers gaining access to proprietary knowledge and potentially using it to build competing products.

“We deploy directly into your VPC or your cloud environment. All the data, that entire data plane is yours to control. We don’t touch it, we don’t see it, we don’t access it.”

Aaron Kao, VP of Marketing, Pinecone

Q: What governance and access controls are built into Nexus for autonomous agent workflows?

Nexus includes access controls and auditability natively, with the ability to cite provenance on every retrieved item and apply access control lists at the document level. Customers can also bring their own models, keeping inference entirely within their cloud boundary. Kao positions these features as the enterprise trust requirements, covering governance, compliance, and audit trails, that Pinecone built into the platform from the start rather than adding later.

“Access controls and auditability built in, the ability to cite where things come from, and access control lists at the documentation level.”

Aaron Kao, VP of Marketing, Pinecone

Q: Where do the cost savings from Nexus appear first: tokens, infrastructure, or engineering?

Kao identifies tokens as the most immediate and visible savings, with 90% reduction in token costs as the headline figure. Engineering costs follow: the total cost of ownership decreases because teams do not need to build and maintain custom pipelines, research compilation algorithms, or manage agentic retrieval infrastructure internally. Kao acknowledges that large organizations could build an equivalent system but argues that purchasing a fully managed solution yields more savings over time than internal development.

“The TCO of this: you can go build what we built with Nexus, but then it’s the innovation, maintaining the pipelines, researching new techniques. All of that is built in and you get to buy that.”

Aaron Kao, VP of Marketing, Pinecone

Q: What enterprise use cases is Pinecone Nexus already supporting?

Kao describes active use cases across legal research, where large document corpora make precision retrieval critical; customer support, where agents draw on call logs and documentation from a shared knowledge layer; fintech M&A deal rooms, with complex company-specific documentation; insurance, with high document volumes; and patent and IP search. Pinecone also uses Nexus internally to replace BI dashboards by running an agent over BigQuery, Salesforce, HubSpot, and Gong call transcripts, eliminating BI engineering hours for ad-hoc business questions.

“We stop using BI tools, we stop using BI dashboards because we built an internal agent that runs on Nexus. There’s a little stat that shows how many BI engineer hours we save each week and it’s just incredibly high.”

Aaron Kao, VP of Marketing, Pinecone

Q: How does Pinecone Nexus differ from knowledge layer approaches by Microsoft, Databricks, and Snowflake?

Kao positions competitors as either maintaining centrally defined semantic or ontology layers, or federating access to data warehouses and lakes. He argues that centrally defined semantic layers are typically set once and drift out of alignment as business workflows evolve. Pinecone’s differentiator is a manifest-driven model where line-of-business experts, the people the agents work for, curate and tune the knowledge layer themselves rather than relying on a central data team to define it.

“We specifically wanted to be expert and line-of-business driven and allow the business that is building and using the agents to curate and build that knowledge. That is the edge.”

Aaron Kao, VP of Marketing, Pinecone

Resources & Documentation

  • Pinecone, AI knowledge platform including Nexus, vector database, and RAG infrastructure for enterprise agents
  • Pinecone Nexus, bring-your-own-cloud knowledge engine that compiles enterprise data into agent-ready knowledge artifacts

***

👇 Click to Read Full Raw Transcript

Swapnil Bhartiya: If you look at enterprise, they have spent decades building out the data. And almost all of that data was built for people, reports for analysts and dashboards for executives. Now there’s a new user of that data, and that user is not a person. It’s an AI agent. Not just any agent, a lot of AI agents. And the stack that was built earlier was never built for agents. So it doesn’t even know how agents consume knowledge. So what agents do, they piece knowledge together, call after call. In most cases, they guess. And when you guess, the results are of course, garbage in, garbage out. So it’s becoming a very serious problem for enterprises. They have agents, they have LLM, they have data, but nothing is talking to each other. And that’s the problem Pinecone is trying to solve with Nexus. It’s a knowledge engine that turns an enterprise’s proprietary data and workflows into governed agent ready knowledge, which is delivered in a single call. And today we have with us Aaron Kao, VP of Marketing at Pinecone, to just talk about that. First of all, Aaron, it’s great to have you on the show.

Aaron Kao: Hey, swap, it’s great to be here. Thank you.

Swapnil Bhartiya: It’s my pleasure. This is the first time I’m talking to Pinecone, so of course I would love to know a bit about the company itself. Let’s start there and then I’ll ask you the problem that I was trying to explain how close I was to coming to actually explain the problem that you’re trying to solve. But let’s start with the company.

Aaron Kao: Pinecone is the trusted AI knowledge company. We were founded with the mission to make AI knowledgeable. So when we first started, we built what is called the vector database. And initially it was used for search and recommendation systems. And then over time as AI started to take off, we pioneered the technique of retrieval, augmented generation. And then going into now the agent era, we are now built ourselves into a larger platform with things to support agents. And our goal really is how can we power AI applications at scale so it’s accurate, fast and cost effective. So pinecone has over 10,000 customers and a million developers worldwide.

Swapnil Bhartiya: Excellent. Thank you so much. Now let’s talk about Pinecone Nexus, what it is and what core role does it play for enterprise AI agents.

Aaron Kao: So Pinecone Nexus, it is a knowledge engine for agents. So it turns enterprise data and like how you work into trusted knowledge based on your own cloud. So the way you can think of it is like everyone can get the same models, you can buy that but really your moat is your company’s moat, is your knowledge and how people do the work. So when people build agents today, they’ll send their agents to look at a lot of data out there in the company, different databases, a lot of unstructured data, ticket and queue systems. And what happens is you rebuild context on every call and that gets really, really costly. It searches it. So usually around like 85% of like your LLM tokens is spent towards just like searching, searching context across lots and lots of data. It also becomes inaccurate and then, yeah, the costs get expensive. So you might be doing the same types of queries as another person. And over time this all adds up. So we’re. What Nexus does differently is look across your entire enterprise and we compile what we call a knowledge artifact. So we build this knowledge layer and then allows you or the agent. So instead of that context being compiled every time at query time, we do it up front at compile time. And then your agent is able to query that very effectively, very efficiently and move very fast. So compile once, serve many times. And you know, we’re able to see a lot of different benefits in terms of cost, accuracy and latency with that.

Swapnil Bhartiya: Now as you’re explaining earlier, the whole how pine cone came to exist, if you look at the traditional agentic rack, it gets expensive and it also gets slow. Of course, cost is a big topic these days. The whole tokenomics is, you know, not only cost, performance, privacy, a lot of things are there. Hallucination is there as well. How does Nexus bring that cost and latency under control?

Aaron Kao: Yeah, yeah. So as I was saying earlier, you know, compared to agentic rag or just plain pointing Claude code at something, we’re actually on our benchmarks. And the early access customers that we worked with were able to see, you know, 90% fewer tokens, like usage. We’re able to see up to like 30 times faster answers and we’re able to get task completion from like 50 to 60% to like close to like over 90%. And the way we do that is, hey, instead of trying to like your agent, compiling like information together and searching through you know, tons of unstructured documents sitting in your Google Drive, you know, tickets and asana, you know, gong call transcripts sitting in Salesforce and a bunch of data every single time, we upfront compile everything so it’s immediately available. So what we do is we build knowledge artifacts which are like, you know, entity relationships between different things, summaries, where certain resources live. So the agents are like a lot more effective in getting the information back out from the systems rather than spending a ton of time searching. So that obviously will reduce costs and reduce latency just from that.

Swapnil Bhartiya: Can you talk about KnowQL, what it is and how does it work inside the Nexus architecture?

Aaron Kao: KnowQL really is just like how you have SQL for apps to access a data layer. We’re giving agents a structured vocabulary to access the knowledge that they need. So right now, how agents get data, right, they ask long sort of questions and then receive everything back. And there is no structure. So a lot of times you have a question, you have certain filters that you want, you want it to provenance, you want certain confidence, you want the output to be shaped a certain way, you want a budget, you want to apply budget to it. So how do you do all of that? That’s not possible in the current agentic world. So we built KnowQL as a way to have structured queries and structured responses back.

Swapnil Bhartiya: I mean, as you know that every enterprise kind of live in its own proprietary data. And also these days there’s also saying that don’t bring your data to AI, bring AI to the data. Because also you don’t want to send your data out to external providers for whatsoever reason. At the same time, a big general purpose model may not be fine tuned for because data, when you talk about data, data could be structured, unstructured, it could be telemetry. So it’s not like one size fits all. But I don’t want to talk about all of that. I just wanted to focus on two things. One is of course privacy. How does Nexus protect that data that AI is touching and the competitive advantage trade secrets that are locked inside it.

Aaron Kao: Yeah, so one of the trends that we saw over the last year is people are increasingly concerned about where the data lives. They don’t want the frontier model. Companies that own all their data, understand their knowledge and maybe build competitive products to them. So no, I completely agree with you. Privacy is a huge, huge concern. And that is exactly a cue that we took when we built Nexus. So Nexus comes as byoc, bring your own cloud. So we deploy directly into your VPC or your cloud environment. And so all the data, that entire data plane is yours to control. We don’t touch it, we don’t see it, we don’t access it. And this is exactly what our customers have told us that they want. They want the ability to curate and build a knowledge layer and have the agents use it, but they want to keep their own data. So hence the BYOC is the exact model for that.

Swapnil Bhartiya: And since we are talking about data privacy, let’s talk about trust. AI and trust, I think they’re like, you know, oil and water, they don’t mix very well. You cannot fully trust AI. That’s why we build harness all the things. What governance and security features are built into Nexus natively so that also these days a lot of agents are autonomous. There is no human loop in most cases so that we can trust those processes that you know what, yeah, there are proper hardware in place. Let’s do it.

Aaron Kao: Yeah, so a few things. So because it’s byoc, the data once lives all there. Right. And what we also allow customers to do is you can bring your own models if you want to run localized models or whatever models, you can keep everything completely constrained to your own cloud environment. Now within governance, access controls and auditability built in, the ability to cite where things come from and then like access control lists at like the documentation level. But overall our goal is to make this as enterprise ready for it to be trusted. Right. We call ourselves a trusted AI knowledge platform because we are building specifically to the enterprise specifications in terms of what they need for trust, governance, compliance and all that. And that’s exactly how we build Nexus.

Swapnil Bhartiya: And if we talk about cost, where do the savings show up first in token infrastructure engineering time?

Aaron Kao: Yeah, you can probably see it across the board. Definitely. Like the headline number for us is tokens. Right. We’ve seen performance gains of like saving 90% on your token costs compared to just querying within your normal agentic retrieval systems. So I think that’s a big one. I think another one is engineering costs. Right. The TCO of this, you can go, you know, there’s a lot of companies out there with a lot of resources, they can build what we built with Nexus. But then it’s like the innovation that goes into it, to maintain that, to maintain the pipelines to, to research new techniques and algorithms, to compile the knowledge, all of that is built in and you get to buy that. So you’ll eventually see that. Right. I’m sure you can build something like that, but over a longer period of time, you’ll save more money on purchasing a fully managed solution like this.

Swapnil Bhartiya: As you said that Nexus has been out for a bit. Is it possible to share some use cases? You don’t have to name the company, but just to give users an idea of how it is helping teams.

Aaron Kao: We’ve had a few different types of companies use it. So we’ve seen legal, especially legal research. Legal is incredibly document heavy and the corpuses are huge. The ability to find things is incredibly hard. There support as well. So a lot of companies they build support agents and it’s looking across a lot of different call logs, documentation, all that. So being able to have a knowledge layer that is commonly shared is really good for that. We’ve also seen like fintech, like M and A, like deal rooms. So every deal room, very complex set of documentation specific to a company. We’ve also seen stuff also legal around like IP or patent search and so yeah there and then we’ve also worked with you know insurance companies that have well they just have a lot of documents and then another use case that we’ve been exploring is just like how do we bring like structured data with the unstructured data and forming relationships across that. And so there’s some e-commerce use cases where the product manager needs to ask questions about a product line and how that’s doing. And so being able to pull that together, we’ve actually built internally, we stop using BI tools, we stop using BI dashboards because we built a internal agent that runs on Nexus, that runs on our BigQuery, all our call log, Salesforce, HubSpot, all these things are just pulled together agentically into a thing and it just saves. There’s a little stat that it shows like how many BI engineer hours we save each month or each week. And it’s just incredibly high because there’s a lot of business questions that you want to ask that would be almost impossible to build dashboards or build custom SQL queries for and or just correlate like Gong call logs with these things. So we use that internally and we’re able to quickly get an understanding of our business through that. So it’s super exciting what we’re building and just what’s possible here.

Swapnil Bhartiya: Can you talk about who would you consider your competitors in this space and what edge do you have over them?

Aaron Kao: I think competitors in the space now are traditional. So there’s a few different things. There’s people that maintain central ontologies and then there are people that federate access and pull the data. And so you know there’ll be things like Microsoft has something around this, you know, Databricks, Snowflake, all of them have stuff that they built on top of their data warehouses, data lakes and to pull this kind of data, pull unstructured data, structured data. And where we’re better is we allow experts like the line of business people who the agents are working for to build the knowledge. So we have this thing called the manifest so tuning that manifest to be exactly what your knowledge workflow is and a lot of competitors they might centrally define like a semantic layer and that a lot of times is set once and then the business deviates, so we have the ability we specifically wanted to be expert and line-of-business driven and allow the business that is building and using the agents to curate and build that knowledge. So that is the edge.

Swapnil Bhartiya: Thank you so much for breaking all of this down and those who are watching please as Aaron mentioned please go and check pinecone.io and Aaron once again thank you so much but this is very important conversation in this whole AI and data space so I would love to have you back on the show but I really appreciate your time today. Thank you.

Aaron Kao: Swapnil, thank you so much for having me on.

AI Security Patches: Can IT Teams Keep Up? | Simon Ritter, Azul | TFiR

Previous article

Failover Testing Without Risking Production | Alexus Gore, SIOS Technology | TFiR

Next article