Enterprise AI pipelines stall not because of model quality but because unstructured data is fragmented across multiple copied locations, none of which stay coherent with the source. On-premises flash storage prices have risen more than 400% year over year as hyperscalers absorb nearly all available silicon supply, making the old answer of buying more storage or copying data to wherever the GPUs are both economically and architecturally unsustainable. Agentic and multi-agent inferencing workflows demand a single, always-current source of truth; stale copies produce what practitioners call the ghost-in-the-machine effect, where agents reason from outdated context and produce unreliable outputs.
In this interview on TFiR, Brandon Whitelaw, SVP and Head of Products at Qumulo, breaks down how enterprises can architect around data gravity, eliminate redundant data copies, and keep GPU clusters fully utilized across cloud, neocloud, and on-premises environments without re-platforming applications.
Guest: Brandon Whitelaw, SVP and Head of Products at Qumulo
Show: TFiR
Here is what every platform engineer and AI infrastructure architect needs to know.
Technical Deep Dive
Q: How do data copies reduce GPU efficiency in enterprise AI pipelines?
Brandon Whitelaw, SVP and Head of Products at Qumulo, explains that enterprises using multiple clouds and neocloud GPU clusters can end up maintaining between two and eight separate copies of the same unstructured data set, none coherently linked to the source. Each copy must be paid for at current storage prices, which for on-premises flash have risen more than 400% year over year. Whitelaw argues that projecting a single coherent cache from the source to all compute locations simultaneously is both the economically and architecturally correct approach, keeping GPUs fed without redundant storage spend.
“If you go use GPUs that are available across CoreWeave and Lambda and other neoclouds and then across multiple hyperscalers, you could end up with between two and eight different copies of the data, none of which are coherently linked back to the original. And then you got to maintain and pay for that data everywhere. That’s just highly inefficient.”
Brandon Whitelaw, SVP and Head of Products, Qumulo
Q: What is data gravity and why does AI make it worse?
Data gravity is the principle that large unstructured data sets are impractical to move, so compute and applications should travel to the data instead. Whitelaw notes that this model worked when GPU hardware was readily available and affordable, but the AI era has inverted the equation: power availability now dictates where compute lands, from wind turbines to neoclouds to public hyperscalers, and eight of the top ten most powerful AI models are cloud-exclusive. Enterprises cannot simply co-locate GPUs with their on-premises data stores anymore.
“GPUs at the end of the day are power-to-token converters, and the power availability compute has to follow wherever that can be found. Wherever people can find power is where compute will be.”
Brandon Whitelaw, SVP and Head of Products, Qumulo
Q: How has the business value of unstructured data changed with AI?
Whitelaw describes a shift from a linear value curve to a U-curve: data is highly valuable when first created, dips in value as it ages, but regains strategic value as the long tail of historical data becomes the context layer for training custom AI models and for agentic inferencing. Every hospital image ever taken, every mile of autonomous vehicle video, and every factory quality assurance image is now reference context that agents must query in real time. That context has weight, and its volume keeps growing.
“With AI, you now have a U-curve on data, meaning it’s really valuable when you first create it, and then the long tail of all that data becomes the collective knowledge of the organization to train AI models or to reference for inferencing.”
Brandon Whitelaw, SVP and Head of Products, Qumulo
Q: Why does unstructured data need contextual indexing before it reaches an AI model?
Without pre-indexing, an AI model asked to scan billions of files must process every object from scratch on every query, burning massive token counts and incurring repeated compute costs. Whitelaw describes Qumulo’s neural search approach: extracting metadata from file headers, such as well site location, camera resolution, or imaging type, fully indexing that metadata, and exposing it directly to the AI tool so the model can target only the relevant subset of files. This converts unstructured information into semi-structured intelligence without moving the underlying data.
“What we’re talking about is being able to go pull that data out once, pre-refine the data, index it, and then feed the AI models so they don’t have to go back through that raw data itself every time.”
Brandon Whitelaw, SVP and Head of Products, Qumulo
Q: What is driving the flash storage price spike and when does it get better?
Whitelaw traces the spike to a silicon wafer supply chain constraint: producing high-bandwidth memory for GPUs consumes six times the wafer area that the same capacity in NAND flash would require. GPU manufacturers pay a substantially higher premium for that memory, pulling supply away from flash drive production. Hyperscaler committed spend contracts have further concentrated purchasing power: the year-over-year increase in hyperscaler hardware spend from 2025 to 2026 alone exceeds the entire on-premises hardware market. Whitelaw states the supply imbalance worsens over the next two years, with hyperscalers projected to absorb over 90% of silicon output in 2026 and total hyperscaler hardware spend reaching over one trillion dollars in 2027.
“On average a 30-terabyte SSD is about 437% more expensive than it was a year ago. Customers coming up for a normal refresh after three years of owning a storage system are not having a twelve-million-dollar bill, they’re having a sixty-million-dollar bill they didn’t budget for.”
Brandon Whitelaw, SVP and Head of Products, Qumulo
Q: Why do NFS and SMB perform so poorly over a WAN for AI workloads?
Traditional file protocols such as NFS and SMB were designed for local area networks. Their chatty, latency-sensitive handshake patterns mean that even a 100 Gbps WAN link may deliver only 8-10% effective utilization when used to move data to a cloud or neocloud GPU cluster. Despite their age, NFS remains the default interface on nearly every AI pipeline toolchain because most AI frameworks expect a POSIX file interface. The protocol mismatch is a structural bottleneck that cannot be solved by adding more bandwidth.
“If you go just take an NFS or SMB connection and connect it from your data center across to a neocloud or to a public cloud, you’re maybe getting 8 to 10% utilization on that link. The protocols are so chatty and so latency sensitive, they were meant for a local area network, not a wide area network.”
Brandon Whitelaw, SVP and Head of Products, Qumulo
Q: How does Qumulo’s data fabric deliver data to remote GPUs without copying it?
Qumulo’s fabric replaces file protocols on the WAN segment with a streaming transport similar to adaptive video delivery, using 100% of available link capacity rather than the 8-10% achievable with NFS. An intelligent caching layer preemptively stages data in front of the GPU before the model requests it, so the GPU sees all data as local regardless of physical location. Because it is a projection of the source rather than a copy, the data remains fully coherent: any change to the on-premises source is reflected in real time at every remote compute location.
“The GPU sees the data as if it’s all local, but it could be 10 petabytes sitting in your data center a thousand miles away in a different state. It sees it as local and then it intelligently preemptively caches the data over our fabric using 100% of the link instead of only 10%.”
Brandon Whitelaw, SVP and Head of Products, Qumulo
Q: What is the ghost-in-the-machine effect and how does it impact agentic AI?
Whitelaw uses the ghost-in-the-machine label to describe what happens when an agentic inferencing workflow reasons against a stale copy of data that has diverged from the production source. During model training, a snapshot copy is acceptable because the training run is a bounded, one-time process. During agentic inferencing, agents query reference data in real time and must receive current context or their outputs become unreliable. Any architecture that relies on periodic data copies rather than a live coherent projection introduces this failure mode.
“As we shift from a training society on AI to an inferencing society on AI, a copy is no longer a safe way to approach working with data. When you’re having an agentic workflow inference off of reference data sets in real time, it has to be the single source of truth.”
Brandon Whitelaw, SVP and Head of Products, Qumulo
Q: How should enterprises architect shared KV cache for multi-agent workloads?
Current multi-agent deployments often treat each agent as isolated, giving each its own KV cache that holds only its own session context. As agent swarms and multi-agent orchestration become standard, Whitelaw argues that a shared persistent KV cache backed by a network-attached storage layer is the correct architectural pattern, mirroring the shift from individual desktop file stores to shared network file shares that solved the multiple-copies problem for human workers in the 1990s. A shared context store lets any agent pick up where another left off without re-deriving context from raw data.
“You go from individual context memory to shared context memory for agents. That demands a better, more modern approach to the consistent memory state for KV cache. It’s just history repeating itself.”
Brandon Whitelaw, SVP and Head of Products, Qumulo
Q: What architectural steps should enterprises take right now to improve GPU efficiency?
Whitelaw’s primary recommendation is to eliminate data copies as the first intervention: any architecture that stages data into multiple cloud or neocloud environments independently multiplies both storage cost and coherency risk. The second step is to adopt a data fabric that projects a single coherent source to all compute locations simultaneously, enabling workload placement decisions based on GPU availability and model quality rather than data proximity. The third step is to move toward shared persistent KV cache infrastructure so multi-agent workflows can operate from a consistent, up-to-date context without per-agent data redundancy.
“Look for data fabric technology that helps you project a single copy of that data to all those locations simultaneously, so you can use wherever the most available compute is or the best models are without having to copy the data.”
Brandon Whitelaw, SVP and Head of Products, Qumulo
Resources & Documentation
- Qumulo, AI data fabric platform for connecting unstructured data to GPU compute across cloud, neocloud, and on-premises environments
- Azure AI Foundry, Microsoft’s AI services platform referenced as a cloud-exclusive model destination
- AWS Bedrock, Amazon’s managed foundation model service referenced as a target compute environment
- Google Vertex AI, Google’s AI platform referenced as a target compute environment
***
👇 Click to Read Full Raw Transcript
Swapnil Bhartiya: As we all know that data has weight. The more of it you hold, the harder it is to move. And AI keeps adding to that weight. Files, images, video and everything is unstructured. Meanwhile, as we all know, storage hardware is scarce and because of whole AI, the prices keep climbing. So enterprises can no longer ship data to wherever the compute lives. The architecture has to flip. So how do you design for GPU efficiency when the data can’t easily move? That’s the problem Kumlio is trying to solve. And today we have with us Brandon Whitelaw, SVP and head of products at Qumulo to unpack all of that. First of all, Brandon, it’s great to have you on the show.
Brandon Whitelaw: Yeah, thank you so much. Great to be here.
Swapnil Bhartiya: It’s my pleasure. I will talk about this problem area, but before that, I would love for you for our audience to also know a bit about the company. So they do know where we are coming from.
Brandon Whitelaw: Yeah. Qumulo is really the world’s first full stack AI data fabric company. We help customers ultimately connect data to compute, to services in the cloud and to talent anywhere in the world and be able to help persist that data, accelerate data and cache that data to accelerate new AI workloads like never before and ultimately address a lot of the problems you just kind of opened with here. So it was a great connection. Great, appreciate it.
Swapnil Bhartiya: Now let’s talk about data gravity itself for those who may not know what it is and what problem it creates. And why is AI making challenge even worse for organizations?
Brandon Whitelaw: Yeah. So the concept of data having gravity is actually one that for better or for worse, I was a member of helping propagate back in my days with EMC and Dell. And the idea was that unstructured data, in particular unlike smaller databases, was growing so fast and was so big that it was not practical to move the data around. You instead should move the compute, the applications, the users to the data. And that was pretty easy to do when compute was readily available and hardware prices were reasonable. And you know, we didn’t have the GPU hardware crunch that we have today. But the table is really flipped in that it is no longer trivial to move compute. And more importantly, what we’re seeing across the AI landscape is that the primary resource people are going after is actually power. GPUs at the end of the day are power to token converters and the power availability compute has to follow wherever that can be found. Could be a distributed inferencing edge, can be on wind turbines, could be in NEO cloud environments, obviously public cloud data centers, you name it, wherever people can find power is where compute will be. You also have AI models that are not widely available everywhere. In fact, eight of the top 10 most powerful models are cloud only exclusive. So you’re not running them on your data center or in a NEO cloud. How do you get your data effectively to those? And when you’re only using the models as a basic chatbot, this isn’t hard. Obviously if you go into Claude and type a question, it gets to it really fast, no problem. But if you’re doing RAG models, if you’re taking enterprise unstructured data and running them through the models to either train them for your own custom weighted model, or you’re trying to build an agentic workflow off of reference data sets that are large unstructured data, you need to have the flexibility to connect data to wherever those agents are, even if it’s not nicely tucked within the four walls of your data center. And that is what we’re trying to change related to the defying data gravity problem that Qumulo’s data fabric helps to aim for customers to solve.
Swapnil Bhartiya: As you also touched upon that for years, the whole idea was move data to compute. Move data wherever it is. Especially in this era of AI, you should not be doing that. You should bring AI to the data, not take data where the AI. Of course we are not even talking about ingress cost as well here, but because of so many different reasons of data gravity as well. Why is that playbook now changing? What is driving that change?
Brandon Whitelaw: Well, I think there’s a couple things that haven’t changed and then things that have changed. So data is still large unstructured data. You know, for every terabyte of new data that’s created in the world, I think 80 to 90% of that is unstructured data. And for enterprises in particular, unstructured data, the content their users create, sensors create, test vehicles for autonomous driving, you name it. It tends to be a lot of unstructured data again, especially if measured by capacity by terabytes. So the data is growing, it’s growing faster than ever. It’s also become more valuable than ever. Data used to have a pretty typical linear value. It was really valuable right when it was created and then it would go down over time as it aged out. Think about you making PowerPoints or Word documents. It’s really valuable for a certain period of time and then it quickly fades out. Now you might come back to it a year later as a reference or if some project comes back up, but it’s not that valuable anymore. With AI, you now have a U-curve on data, meaning it’s really valuable when you first create it. And then the long tail of all that data becomes the collective knowledge of the organization to train AI models or to reference for inferencing for those models. So if you’re trying to help do something, especially in an agentic inferencing side of AI, not even training, you need an agent to be able to reference back to a core data set for context. That context is your enterprise unstructured data. It’s every image a hospital has ever taken to help look for lung cancer. It’s every mile of video recorded for autonomous vehicles. It’s every quality assurance image taken at a factory for new products. That context has value and that has weight. What does that weight mean? Again, it’s just large in capacity. Now, again, when compute was wildly available, when anyone could walk off the street and say, hey, I need a server with some GPUs, then it was pretty easy to move those GPUs to where the data was. And the GPUs didn’t consume that much power relative to today. I mean, they did compared to other things, but not till today. You have circumstances where in some customers I’ve dealt with, especially in Europe with limited data center space and power, they can’t put a single GPU in a rack of the newest models. It’s too much power. And so now we have a circumstance where people are saying, well, geez, I need access to GPUs, I can’t put them on my data center. I don’t have the power, I can’t build a new data center. I mean, if you want to build a 20 megawatt data center, you have a two year waiting list to break ground and that’s looking at a six year lead time on backup generators and other things. And so everyone’s basically saying, how do I connect my data to wherever I can find compute, which is wherever I can find power and wherever the best models are for the task at hand. And so what we’re trying to do is help customers defy data gravity and being able to leave data where it is, but project it intelligently into wherever those models in compute are so they can move their business forward.
Swapnil Bhartiya: Also another problem is that much of this data that we’re talking about is also unstructured and it’s also keep growing in size. What does that growth mean for enterprises? Storage of state?
Brandon Whitelaw: Well, I think it means a couple things. One is they have to work on projects to move from having information to intelligence. Today most data is just information, but they don’t truly understand what they have nor its potential value. Now that used to be just how can I make a better search bot for my employees? Or maybe some OCR indexing of images so I can search for if the Social Security number is on an x-ray. But now people are wanting to be able to empower these AI models and the AI models need more context to understand what data to train against. So what we’re working on with like for example our neural search products is not just fully indexing the file system to feed a rapid parquet file index into an AI model, but contextual indexing, meaning let me extract additional metadata information from the file headers. This could be the well site location, the type of well, this could be the type of camera that was used and the resolution to take that video, you name it by industry, there’s very different use cases. Take that information out of the file itself, fully index it, and then allow an AI tool to have direct access to that so it can more rapidly and cost effectively find the right data to do what it needs for the job. You know, if you were to go pull up Claude today and go look at a hospital medical imaging system, you would have sometimes tens of billions of images. And if you were to say, hey Claude, let me go look through this entire repository and tell me if any of these images have Social Security numbers that aren’t masked appropriately and therefore violating HIPAA compliance. Well, it’s going to go through and try to look at every single image and burn a bazillion tokens and millions of dollars trying to do that. And every time you ask that question, it’s going to do that again. And so what we’re talking about is being able to say how can I go pull that data out once, pre-refine the data, index it and then feed the AI models so they don’t have to go back through that raw data itself every time. I’ve effectively taken unstructured data and made a semi-structured index that can feed those AI models and move it from information to intelligence.
Swapnil Bhartiya: Now. Very well said. Thank you, thank you so much. Now in addition to that, another problem is of course the hardware cost is also becoming a big challenge. So it’s not just where your data is, how you are giving it to AI, it’s also about earlier, you know, data storage was so cheap, I mean that, you know, everybody was just capture everything, we’ll figure out what value to extract from it. But now because of the cost, you have to rethink the whole storage economics as well right now. Which may be a good thing or a bad thing, depending on earlier we used to say, hey, storage is so cheap, just go buy. No, not anymore. RAM is not cheap anymore. So how are they rethinking the whole storage economics or what is your advice to them?
Brandon Whitelaw: You know, the cost of status quo has never been higher. So it used to be the way you solve the storage problems, you just buy more storage. Like hey, I have more data, just let me buy some more storage, nice and easy, no problem. I don’t have to reconsider my architecture where it lands. And actually, in fact it was so cost effective at one point that everyone was saying, I’m just going to move everything to all flash, even for tens of petabytes of unstructured data, even for just basic home directories and file shares where they don’t need the performance of all flash. But I’m going to do it just simply because it’s cost effective enough. And you know, who doesn’t want a little bit more performance? And maybe the drives don’t fail as often. And so other things like this, there was perceptions that made it so. You know, think about it this way. If all of a sudden Ferraris were the same price as Honda Civics, everyone would start buying Ferraris, right? Or if it got close enough, let’s put it that way. Now what’s changed is that we’ve had a supply chain domino effect occur. So as we’ve had to build more GPUs, one of the key components of a GPU is high bandwidth memory. Now high bandwidth memory is actually made from the same silicon wafers that makes DRAM and NAND flash, you know, SSD and NVMe drives. But the problem is it’s a 6 to 1 ratio. So you make one-sixth the amount of high bandwidth memory out of the same wafer that you would normally make for NAND flash. What would normally be a maybe six terabytes of flash now becomes just one terabyte of high bandwidth memory. And the GPU developers, you know, Nvidia, AMD, etc., they will pay a substantially higher premium for that high bandwidth memory than a normal flash drive manufacturer because it’s attached to a really expensive thing. That GPU processor is actually ridiculously more expensive than even the high bandwidth memory next to it. So in contrast, it looks pretty affordable. Whereas when you’re buying all flash storage systems, that’s a completely different economic equation about what you’re willing to pay for that and also who’s buying how much has changed. So you know, back in 2021, the amount of hardware consumed by the hyperscalers, the five large public cloud providers, plus Meta plus compared to what we call the big seven enterprise vendors, this is your Dell, HPE, IBM, Lenovo, etc., all the enterprise storage and server manufacturers, they were about the same, about 100 to 120 billion dollars a year back in 2021. Well if you fast forward to last year, the hyperscalers were at 450 billion and the big seven enterprise hardware vendors on prem were only at about 150 billion. So all of a sudden went to about equal to three times more spend by the cloud than the big seven enterprise vendors on prem. And then this year it went just crazy. So this year we’re sitting at about 750 to 800 billion compared, in fact just the year over year change from last year to this year by hyperscaler spend on hardware is more than the entire on-prem hardware business collectively across all hardware manufacturers, all server manufacturers, all storage manufacturers. So the hyperscalers have started buying up all the supply. Why? Well, because they found that customers were starting to pick which cloud to put their workload in based on availability of resources, not cloud preferences or high value services. You know, it used to be you had an AWS customer and they preferred everything in AWS or Azure or GCP. And now you have customers saying yeah, I prefer one, but because I can’t get availability to GPU or compute, I’m going to move this workload to another one that has it available. So they started in this battle on who’s got compute availability. And they said okay, well how do we get more compute availability? Well, I have to get first in line for buying more of it. All the big hyperscalers said not only will I guarantee to spend a certain amount, I will guarantee over three years to spend a certain amount with escalating contracts. So while we’re talking about 700 to 800 billion this year, the estimate for next year is over a trillion and then 1.3 trillion the year after that. And the amount of supply coming on market from all the new fabs and otherwise, actually the problem gets worse over the next two years, not better. So we still have a lot of runway where the demand wildly outstrips supply. This year about 70% of silicon is going to hyperscalers. Next year it’s estimated to go over 90%. And because of that the on-prem price for flash has quadrupled year over year. On average a 30-terabyte SSD is about 437% more expensive than it was a year ago. So you have customers that were coming up for a normal refresh after three years of owning a storage system and not having a 12-million-dollar bill, having a 60-million-dollar bill they didn’t budget for. And even the hard drive side of things are up about 60% year over year because as people are going, oh my gosh, flash is too expensive, now I’m going to buy more hard drive based storage systems, that has driven up the demand. Western Digital has sold out all capacity for the next two years. Same with others. So again, who’s buying most of that? Hyperscalers, right? All the big object systems in the cloud. So because of this change of dynamic, customers are having to start to rethink how they architect the modern data landscape away from just everything in their data center. How do I get access to that storage capacity in the cloud while being as least disruptive as possible to what I’m doing with my clients and users and access. And again our infinite capacity solution, helping extend on-prem systems in the cloud, is allowing them to do that so they can get a bottomless bucket of storage without having to re-platform and rewrite applications.
Swapnil Bhartiya: If you forget about all those things and just focus on AI and the success, the success also depends on fast access to this data. Talk about the impact it makes when data has to travel a lot on the performance you get from AI.
Brandon Whitelaw: So I think one thing that’s really been having an impact on organizations is not just the amount of data but the fact that where they need to get that data to, hyperscalers’ AI as a service platforms like Azure Foundry, AWS Bedrock, Google Vertex, or get access to the models that are cloud exclusive, or to get access to GPUs in NEO clouds because they have the power contracts. This means traversing data out of your network into other systems over a WAN usually. And traditional file protocols were really not made for the WAN. So if you go just take an NFS or SMB connection and connect it from your data center across to a NEO cloud or to a public cloud, you’re maybe getting 8 to 10% utilization on that link. You might have a 100 gig link, you may struggle to get more than even a tenth of that link to be used. Because the protocols are so chatty and they’re so latency sensitive, they were meant for a local area network, not a wide area network. So because of this, even if you had the bandwidth, you can’t get the performance out of it through the protocols. And for better or for worse, most AI tools still require a POSIX file interface, mostly NFS, as crazy as that sounds, because that’s the oldest file protocol of them all. That is the default on almost every AI pipeline toolchain. And so what we have tried to do is help customers transition where from the client to the data, it’s still standard protocols. And then over our data fabric, it doesn’t use file protocols. In fact, it streams data more akin to like a Netflix stream, and then puts that right in front of the GPU, intelligently cached so it can use that data in place. So the GPU sees the data as if it’s all local, but it could be 10 petabytes sitting in your data center a thousand miles away in a different state. It sees it as local and then it intelligently preemptively caches the data over our fabric using 100% of the link instead of only 10%. So you get more of the bandwidth you have available and you also get the data there before it’s even needed. Because we have again an intelligent caching model that is very particularly attuned for these workloads to make it so you can keep the GPUs very happy and full, even though you’re not having to pre-stage or copy the data there. And by the way, that you can almost think of it as like kind of an anti-copy technology. You know, the traditional approach today is like, well, I got to go train on this data or use this data for this AI model, I’ll go copy the data from my data center into a cloud, into a NEO cloud. The challenge with copying data, of course, is that A, paying for things twice when hardware is the price it is, is not a great solution. And then B, the moment you make a copy of data, it’s no longer coherent with the original. So if you make changes on this production data, it’s out of sync with the copy. And now you’re agentically inferencing on a stale data set. That’s not going to work well. So part of our value with our data fabric, our AI data fabric, is making it so that you have strong consistency and coherency to the data. It’s never a copy, it’s a projection of the original. And it’s always maintained and up to date so that the moment you make any changes to the source, it’s reflecting real time into those AI workloads, wherever they may be.
Swapnil Bhartiya: Now let’s talk about efficiency. How should, because these, we have no control over hardware availability, we have no control over hardware pricing. But there are certain things that we can do. So how should enterprises architect for GPU efficiency, what techniques they can use to make it more efficient and what does it look like in practice?
Brandon Whitelaw: Yeah, I think one of the first things I would do is look at anything that can help you avoid copying data. Because ultimately at the end of the day, if you want to use GPUs that are available across CoreWeave and Lambda and other NEO clouds and then across multiple hyperscalers, each have some of their own unique AI models and capabilities. You could end up with between two and eight different copies of the data, none of which are coherently linked back to the original. And then you got to maintain and pay for that data everywhere. That’s just highly inefficient. So instead, looking for data fabric technology like our own within Qumulo that helps you project a single copy of that data to all those locations simultaneously. So you can use wherever the most available compute is or the best models are without having to copy the data. It also means that especially as we shift from a training society on AI to an inferencing society on AI, this means that a copy is no longer a safe way to approach working with data. It’s fine when you’re training, but when you’re having an agentic workflow inference off of reference data sets in real time, it has to be the single source of truth. Otherwise you get what’s called a ghost in the machine effect. And so again, projecting a strong coherent cache across from the source of data is the more efficient way. And the last bit also comes down to KV cache. We’ve kind of treated agents in isolation. Each agent has its own KV cache. It has the session memory of everything you’ve been talking about with them. But now as we get swarming effects of agents and multi-agent workloads, you want to actually have a shared persistent session memory where any agent can reference the same context that any other agent was working on and pick up where they were and move on with it. And so this more and more forces architects to look at, how do I, you know, kind of repeat history. Back in the day, everyone had their own individual machine at work and everything that they created was on their own desktop and it was available just to yourself. And if you wanted to share that data, what did you do? You made a copy and emailed it around, right? And then there were eight copies of the same PowerPoint. Well, what solved that is a network attached storage system where everyone could put it on a file share and everyone could see it in one place and there was one version of it. Well, it’s no different for agents. In fact, it’s just history repeating itself. You go from individual context memory to shared context memory for agents. And that demands a better, more modern approach to the consistent memory state for KV cache. And systems like ours are meant to help address that challenge for customers.
Swapnil Bhartiya: Brandon, thank you so much for walking us through. Of course, this I really appreciate your time today and of course this is a topic which is also one of the pressing problems right now. So I would love to have you back on the show to understand what work you folks are doing and how you folks can help enterprises. But I really appreciate it. Brandon, thank you.
Brandon Whitelaw: Great being here. Look forward to the next time. Appreciate it.





