AI Infrastructure

How to Make Enterprise Data Searchable for AI Agents | Michel Tricot, Airbyte | TFiR

0

AI agents retrieve only what they can find. If enterprise data is stored without indexing, search capability, or encoded company context, agents return generic or inaccurate answers regardless of model quality. The problem is not at the model layer. It is at the data layer that was designed for human analysts, not autonomous systems. As organizations move toward agentic workflows, the infrastructure assumption that storage format is the primary concern is proving to be the wrong starting point.

In this interview on TFiR, Michel Tricot, Founder and CEO at Airbyte, breaks down why searchability is now the defining constraint in AI data infrastructure, how the semantic layer concept is re-emerging for agentic pipelines, and what an intelligent, self-improving data layer will look like as agents become central to enterprise operations.

Guest: Michel Tricot, Founder and CEO at Airbyte
Show: TFiR

Here is what every data engineer and AI infrastructure team needs to know.

Technical Deep Dive

Q: Should organizations change how they store data to prepare for AI, or is current storage good enough?

Michel Tricot, Founder and CEO at Airbyte, argues that the storage format itself is less critical than whether that data can be searched effectively. The real shift is not about changing how data is generated or stored at a structural level. It is about ensuring that whatever is stored is indexed in a way that surfaces the most accurate and valuable information first when an agent queries it.

“It’s not so much about how you store it; it’s more about how you store it so that it’s searchable and indexed in a way that always provides the most accurate and valuable information first.” — Michel Tricot, Founder and CEO, Airbyte

Q: Where does search fit into the evolution of AI data infrastructure?

Tricot identifies search as the most important capability in the next evolution of data infrastructure. Agents require a very strong ability to search data and to discover what is available to them before they can act reliably. The infrastructure investment priority is shifting from pipeline throughput and storage cost toward indexing quality and search precision.

“The thing that is the most important is the ability to search. An agent will need to have a very, very strong ability to search data to discover what is available to it.” — Michel Tricot, Founder and CEO, Airbyte

Q: What is the actual role of vector databases in AI data infrastructure?

Tricot frames vector databases as a search mechanism for unstructured data rather than a fundamentally new storage paradigm. The hype around vector databases emerged because organizations suddenly recognized that unstructured data could be made useful, and vector search provided a way to query it. At the infrastructure level, vector databases are one implementation of the broader searchability requirement that agentic systems place on data.

“Vector databases is an ability to search around unstructured data. Because suddenly we realize, oh yeah, we can do something with that unstructured data.” — Michel Tricot, Founder and CEO, Airbyte

Q: Is a dedicated data layer between AI systems and enterprise applications actually coming?

Tricot is direct: this layer is coming, and sooner rather than later. While model architectures will continue to change, one constant remains. Making models relevant to a specific organization requires injecting that organization’s data. The intelligent data layer will bring additional logic to guide the model and to help it build connections between different records and silos, not just retrieve raw documents.

“Something that will not change is the fact that to make these models relevant, you need to inject data.” — Michel Tricot, Founder and CEO, Airbyte

Q: How does the semantic layer concept from data warehousing apply to agentic AI pipelines?

Tricot draws a direct parallel between the semantic layer that data warehouse teams built to enforce shared definitions and the contextual encoding that AI pipelines will require. The semantic layer for warehouses encoded a type of memory into the system so that everyone was working from the same language. The same pattern is now emerging for AI, where company-specific knowledge must be encoded as an overlay on top of the raw data so that agents understand context, not just content.

“A few years ago, everyone was talking about the semantic layer for data warehouses. That was a way of encoding some information, some type of memory into the warehouse to make sure that everyone is talking the same language. I think that’s going to be the same pattern that’s going to apply to data.” — Michel Tricot, Founder and CEO, Airbyte

Q: What role does memory play in making AI agents useful inside a specific organization?

Tricot uses a concrete example: Airbyte operates on a fiscal year that starts in February. An agent with no access to that context cannot correctly interpret financial data tied to that calendar. Memory, in this context, is the mechanism by which company-specific knowledge gets encoded as an overlay on top of data so agents can interpret it accurately. The goal is for this encoding to happen automatically over time, improving based on the relevance of answers the model has already provided.

“How do you make sure that over time, all this company knowledge gets encoded as an overlay of the data and ideally automatically?” — Michel Tricot, Founder and CEO, Airbyte

Q: Will the intelligent data layer be human-guided or self-improving?

Tricot expects both. The initial connections and context may be guided by humans, but the system will also self-improve based on feedback signals such as the relevance of answers the model returns. This continuous enrichment of the connection layer between records and silos is what makes memory in AI systems more than a static knowledge base. It becomes an evolving representation of how the organization actually works.

“Yes, it might be guided by humans, but there will always also be a self improvement based on how the relevance of the answer has been provided by the model to just enrich this connection.” — Michel Tricot, Founder and CEO, Airbyte

Resources & Documentation

  • Airbyte, open-source data integration platform for moving data from sources into warehouses, lakes, and AI pipelines
  • Airbyte on GitHub, source code, connectors, and contribution documentation for the Airbyte platform

***

👇 Click to Read Full Raw Transcript

Swapnil Bhartiya: As you’re also talking about organizations, they have semantic data. They have to look at it because of AI. Should they change the way they store data or because of AI, they really don’t have to worry too much because AI can really filter, snip through all the data and pick the data that it needs. What I’m trying to ask you is that should organization change their approach to how they generate, create, store data or it really doesn’t matter, they could continue as they are doing, I would say

Michel Tricot: today the thing that is the most important is the ability to search. So if I’m thinking about new evolution of data infrastructure, I think a lot more is going to be put on search. You know, we had this whole hype around vector databases. At the end of the day, vector databases is an ability to search around unstructured data. Because suddenly we realize, oh yeah, we can do something with that unstructured data. But at the end of the day, an agent will need to have a very, very strong ability to search data to discover what is available to it. And that’s where I see a lot of the infrastructure moving toward is it’s not so much about like how do you store it, it’s more like how do you store it so that it is searchable, so that it is indexed in a way that always provides the most accurate and the most valuable information first.

Swapnil Bhartiya: Where do you see this is going? Are we heading towards new data layer that sits specifically between AI systems and enterprise application or this is too early, we’ll see how things will pan out?

Michel Tricot: No, there is definitely something that’s going to happen and sooner rather than later, the models might change. The architectures of the models will change over time. That is a certainty. Something that will not change is the fact that to make these models relevant, you need to inject data. So what I think is going to happen on this data layer is how do you bring more intelligence to guide the model into or to get the model to actually populate more information about what kind of connection is being able to do between different records, between different silos. And that to me is a few years ago, everyone was talking about the semantic layer for data warehouses. That was a way of encoding some information in a way some type of memory into the warehouse to make sure that everyone is talking the same language. And I think that’s going to be the same pattern that’s going to apply to data, except that yes, it might be guided by humans, but there will always also be like a self improvement based on how EVAs are going the relevance of the answer that has been provided by the model to just enrich this connection. And that’s where we’re talking about memory a lot. This is where memory is going to become very important. Airbytes has a fiscal year that starts in February. Well, if the agent doesn’t know, it cannot understand the data. So how do you make sure that over time, all this company knowledge gets encoded as an overlay of the data and ideally automatically?

Data Maturity Is the Real Bottleneck for Enterprise AI Agents | Mario Moscatiello, Airbyte | TFiR

Previous article