Enterprises building on large language models are discovering a structural problem: every time proprietary business context is sent to a frontier AI provider, that context leaves the organization’s control. For regulated industries this has always been a hard legal boundary. For everyone else, the competitive cost is now becoming visible.
In this interview on TFiR, Mario Moscatiello, VP Growth at Airbyte, covers how Airbyte’s open source, hybrid deployment architecture enables organizations to provide governed data access to AI agents without that data ever leaving their own environment.
Guest: Mario Moscatiello, VP Growth at Airbyte
Show: TFiR
Here is what every data engineer and platform architect responsible for AI infrastructure needs to know.
Technical Deep Dive
Q: Why is AI data sovereignty becoming a priority for organizations outside regulated industries?
Mario Moscatiello, VP Growth at Airbyte, explains that first-party data has always been the core competitive asset for any company, but the shift to AI has made the risk of data leaving the organization newly visible. When companies send proprietary context to frontier AI providers, they are effectively giving those providers their strategic intelligence and then paying to access derived outputs from it. Non-regulated software companies that previously had few data residency concerns are now recognizing this as a direct competitive risk.
“Your first-party data is your gold, and as much as you can, you should always strive for data not to leave your environment.” — Mario Moscatiello, VP Growth, Airbyte
Q: How have regulated industries such as finance, insurance, and healthcare handled AI data residency requirements?
Moscatiello notes that regulated industries have not had a choice on this question. Legal mandates have required that data for finance, insurance, and healthcare organizations cannot leave their environment, so on-premises and hybrid data infrastructure has been standard practice in those sectors for years. This compliance-driven approach has effectively made regulated industries the reference model that non-regulated companies are now choosing to follow voluntarily.
“If you’re a finance organization or insurance or healthcare, your data by law cannot leave your environment.” — Mario Moscatiello, VP Growth, Airbyte
Q: What does “renting back your intelligence” mean in the context of frontier AI providers?
Moscatiello describes a pattern where organizations provide all of their business context to frontier AI labs and then consume AI-generated outputs as a paid service. The problem is that the intelligence driving those outputs was derived from the organization’s own proprietary data. The company has transferred its competitive knowledge to an external provider and is then purchasing access to a version of it, rather than retaining that value internally.
“We’re giving frontier labs all of our context and then we’re renting back our intelligence in that sense.” — Mario Moscatiello, VP Growth, Airbyte
Q: How does Airbyte support on-premises and hybrid deployments for AI data pipelines?
Moscatiello explains that Airbyte has supported on-premises deployment for years as a direct consequence of being open source. Organizations can self-host Airbyte within their own infrastructure, meaning data movement for AI pipelines occurs entirely within the organization’s environment. This hybrid deployment capability is now the primary reason non-regulated companies are evaluating Airbyte as an alternative to cloud-only data integration services.
“We’ve been deploying on prem for years because we are open source.” — Mario Moscatiello, VP Growth, Airbyte
Q: What does governed data access for agentic AI look like in practice with Airbyte?
Moscatiello describes inbound requests from companies asking specifically for a system that can supply data to AI agents without that data leaving the organization’s environment. The goal Airbyte is working toward is providing governed data access that remains fully within organizational control, so that the proprietary context powering AI agents is never exposed to external infrastructure. This positions Airbyte as the data access layer between internal data assets and agentic AI workloads.
“How do we provide governed data access that is fully within an organization’s control so that companies are not really selling their secret sauce.” — Mario Moscatiello, VP Growth, Airbyte
Q: How does geopolitical uncertainty and EU AI regulation affect organizational decisions about AI data infrastructure?
Moscatiello acknowledges that the broader environment of geopolitical uncertainty and incoming AI-specific legislation, particularly in the EU, is accelerating demand for data sovereignty solutions. Organizations and countries are increasingly prioritizing control over their AI infrastructure regardless of whether a specific regulation currently applies to them. The combination of regulatory pressure and competitive risk is driving adoption of on-premises and hybrid data pipeline architectures across sectors that previously had no formal data residency requirements.
“A lot of companies and countries want AI sovereignty, and that is what we want to help with.” — Mario Moscatiello, VP Growth, Airbyte
Resources & Documentation
- Airbyte, open source data integration platform with on-premises and hybrid deployment support for AI and analytics pipelines
- Airbyte on GitHub, self-hosted deployment repository for running Airbyte within your own environment
***
👇 Click to Read Full Raw Transcript
Swapnil Bhartiya: Now there’s a lot of geopolitical crisis going on and decisions are being made which may or may not be good for AI. A lot of anti AI fuddies also going on. So we are living in an era where there are a lot of advantages and then also there is a lot of skepticism also going on there. Regardless what we are seeing is that more and more organizations, more and more companies, more, more and more countries, they do want AI sovereignty. How does Airbyte facilitate that? Because we are also, you folks also work in regulated industries, compliance centering industries and also when you throw the whole of course in eu, a lot of AI related laws are coming into force as well. So talk about how is Airbyte helping organizations countries to take control of their AI inside out.
Mario Moscatiello: Yeah, I think something we’ve been saying a lot for the past two years at Airbyte especially is that when you’re working in these sensitive applications, but also when you’re not like now it’s like anybody, your first party data, it’s your gold essentially as a company and as much as you can, you should always strive for data not to leave your environment essentially. And now this has been true for regulated industries simply because they didn’t have a choice. If you’re a finance organization or insurance or healthcare, well, you have no choice. Your data by law cannot leave your environment. But when you’re a software company like Cloud Native, maybe that matters a bit less. Until now because what we’re seeing is that a lot of companies are saying yes, we’re giving Frontier Labs and all of these providers all of our context and then we’re kind of like renting back our intelligence in that sense. And so like we, you know, we, we have hybrid deployments, we’ve been deploying on plan on prem for, for years because we are open source. And so what we’re seeing is a lot of companies even that are from non regulated industries are coming to us and saying hey, like do you have a system for which I can give data to my agents without that data leaving my environment? And that’s sort of what we’re seeing today and that’s what we want to help with and that’s what we want to like you know, focus on is how do we provide governed data access that is fully within an organization control so that you know, companies are not really selling their secret sauce.





