Cloud NativeAI Infrastructure

Telemetry Costs Are Unsustainable. The Bring-Your-Own-Cloud Fix | Gabriel-James Safar, Tsuga | TFiR

0

Telemetry volumes are growing faster than observability budgets, and the standard response, sampling more aggressively, guarantees you will not have the data you need when an incident hits. At the same time, AI agents and LLMs are introducing new classes of service that traditional platforms were never instrumented to handle, including sensitive data leakage through telemetry pipelines and unpredictable inference costs. Closed-source collectors have locked enterprises into vendor ecosystems for years, and the cost of switching feels prohibitive precisely because the data and the pipeline are both owned by the vendor.

In this interview on TFiR, Gabriel-James Safar, CEO at Tsuga, covers the structural failures of legacy observability platforms, the bring-your-own-cloud architecture Tsuga uses to return data sovereignty to enterprises, and the organizational and technical approach required to migrate large-scale systems without operational disruption.

Guest: Gabriel-James Safar, CEO at Tsuga
Show: TFiR

Here is what every platform engineer, observability lead, and infrastructure architect needs to know.

Technical Deep Dive

Q: What is Tsuga and what problem did it set out to solve?

Gabriel-James Safar, CEO at Tsuga, co-founded the company after leaving Datadog, where he and co-founder Sebastian had led a large suite of products following an acquisition. After speaking with large enterprises, Safar identified that the Datadog model, while strong for many customers, was not keeping pace with the structural shifts affecting the largest organizations. Tsuga was built specifically for enterprises dealing with sovereignty requirements, governance complexity, and observability at very large scale. The name comes from a family of pine trees native to Japan and the Pacific Northwest, chosen partly because of Safar’s personal connection to Japan through his wife.

“The Datadog model was great in a variety of ways. It was the best product for many customers. But at the same time it was not following a variety of shifts that we were seeing with notably larger organizations.” — Gabriel-James Safar, CEO, Tsuga

Q: How are AI agents and LLMs changing what enterprises need from observability platforms?

LLMs and AI agents represent a new category of service that must be monitored in production for performance, cost, and correctness, just like any other service. Because AI-assisted development accelerates deployment frequency significantly, teams need to retain far more telemetry data to detect regressions introduced by the higher volume of code reaching production. A third structural issue is data sensitivity: LLM traces can contain personal health information or other sensitive customer data that should not be forwarded to a third-party observability vendor.

“If you’re a company with healthcare and one of your customers is telling personal information about their health, that information will be in the telemetry of your LLMs. Is it okay to have that information sent to a third party? Probably not.” — Gabriel-James Safar, CEO, Tsuga

Q: How did observability evolve before AI, and did AI change that trajectory or accelerate it?

The move away from closed-source telemetry collectors had already been a top priority for large enterprises for roughly five years before AI became a dominant concern. Closed-source agents created deep vendor lock-in by embedding proprietary collection logic directly into application infrastructure. AI has made the transition to open-source collectors easier and faster, but it has not made it instantaneous. The more significant AI-driven shift is that telemetry is now recognized as business data, and enterprises want the ability to feed it directly into their own BI tools and AI models, which requires owning the data outright.

“Getting rid of these locked closed-source agents has been a top priority project for the past five years in many enterprises. What AI is bringing to that is that the transition is much simpler, much easier using AI to do the transition.” — Gabriel-James Safar, CEO, Tsuga

Q: What is driving enterprise demand for data sovereignty in observability?

Regulatory pressure is accelerating across jurisdictions including Canada, Brazil, Japan, India, Australia, and the GCC countries such as Saudi Arabia and the UAE, all of which impose requirements on where customer data is stored. Beyond regulation, Safar frames sovereignty as a company-level concern as much as a country-level one: if a vendor holds your telemetry data, that vendor could in principle use it to benefit your competitors. For large enterprises, owning the data pipeline end to end is a strategic priority, not only a compliance one.

“The notion of sovereignty, I don’t think we should hear that only at country level. We should also see that at company level, notably for companies that are important. If you’re a big company, you are a sovereign entity.” — Gabriel-James Safar, CEO, Tsuga

Q: How does Tsuga’s bring-your-own-cloud architecture work and what does it give enterprises?

Tsuga is compatible with all open-source data collectors including OpenTelemetry, and the entire observability stack, storage, processing, and the core product, runs inside the customer’s own cloud environment. This means telemetry data never leaves the enterprise boundary, there is no infrastructure tax paid to Tsuga, and the data can be fed directly into the customer’s own BI tools or AI systems. Tsuga also supports a bring-your-own-agent model for AI, meaning enterprises apply their own approved AI systems on top of the telemetry rather than relying on a vendor-chosen model.

“The data is yours. You don’t pay a tax to use it. And on top of it, even when we look at AI, we work with the bring-your-own-agent approach. So you can choose your AI systems and run it on top of the telemetry.” — Gabriel-James Safar, CEO, Tsuga

Q: Why is aggressive sampling a flawed response to rising telemetry costs?

Traditional platforms have responded to volume growth by encouraging ever-higher sampling rates, with some enterprises retaining as little as 0.1 percent of telemetry. The problem is that you cannot predict in advance which traces or metrics you will need for troubleshooting or analysis, so the more aggressively you sample, the higher the probability that the specific data required during an incident is gone. This creates compounding operational cost: engineering teams spend significant time managing sampling pipelines while still being left without adequate data when it matters.

“The more you sample, the more you create operational problems for your teams because they need to spend a lot of time reducing the volumes, and the more likely it is that you won’t have the right data when you need it to do analysis or troubleshooting.” — Gabriel-James Safar, CEO, Tsuga

Q: How does Tsuga’s pricing model differ from traditional observability vendors?

Because Tsuga runs inside the customer’s own cloud rather than on Tsuga-owned infrastructure, the company does not need to apply a large infrastructure markup to its pricing. Traditional vendors typically price at five times infrastructure cost to achieve high gross margins, which structurally limits how much data customers can afford to retain. Tsuga’s model is designed so that customers can retain roughly ten times more data than they could with a comparable traditional platform, and because the data sits in their own cloud environment, they may also be able to negotiate better rates directly with their cloud provider.

“The goal is to flip that and to say we are going to make it work so that even if you have a huge amount of data, the pricing allows you to keep maybe ten times more data than what you could do with another system.” — Gabriel-James Safar, CEO, Tsuga

Q: What is Tsuga’s position on open source versus proprietary components in observability?

Safar draws a clear boundary: the data collectors and the storage format must be open source so that enterprises are never locked into Tsuga’s ecosystem and can continue to use their data in Databricks, BigQuery, Athena, or other tools if they leave. The central product layer is currently proprietary because delivering an opinionated, high-quality product at that level is difficult to do in the open. However, Safar acknowledges that the open boundaries on both sides create real competitive pressure on Tsuga to deliver sufficient value, because a customer can exit while keeping their entire data pipeline intact.

“If we don’t deliver enough value, our customers have open-source data collectors at the entrance, open-source data at the exit, and so they can get rid of us and keep their data flow end to end even without us.” — Gabriel-James Safar, CEO, Tsuga

Q: What is the biggest blocker when enterprises try to migrate away from legacy observability platforms?

The primary blocker is people, not technology. Observability tools are embedded in the workflows of a very large number of teams across an organization, and any migration disrupts those workflows. Tsuga addresses this by working with the central platform or observability team to show how the transition makes their day-to-day work easier, backed by governance tools that help enforce telemetry quality and observability standards company-wide. The secondary blocker is the data collection layer itself, particularly for organizations that have been using closed-source collectors and need to re-architect the pipeline using OpenTelemetry agents.

“The biggest blocker are the people. How can we empower the people so that their life is going to be easier? That’s the biggest thing that needs to be identified with the central team.” — Gabriel-James Safar, CEO, Tsuga

Q: How does Tsuga support multi-region and multi-cluster data residency requirements?

Tsuga supports multi-cluster deployments, allowing an enterprise to keep separate partitions of telemetry data in different geographic regions such as the US, Europe, and Brazil, while presenting a single unified interface to the operator. Forward-deployed engineers from Tsuga help customers design the data flow architecture, determine what telemetry should be routed where, identify pipeline bottlenecks, and implement the configuration. This service layer is central to how Tsuga positions itself: as a software and service company rather than a pure SaaS platform.

“We like to say that we are not a SaaS, but we are a SaaS. We are a SaaS in the sense that we are a software and a service.” — Gabriel-James Safar, CEO, Tsuga

Q: How is Tsuga investing its Series A funding?

Tsuga raised $30 million and is allocating the capital across three areas. Product development is the first priority, as there is significant roadmap ahead even given the founding team’s deep domain experience. Sales expansion is the second priority, with the goal of growing coverage beyond the current team in France, Germany, the UAE, the US, and the UK by hiring salespeople who specialize in complex enterprise deals and pairing them with forward-deployed engineers. Marketing investment to support the enterprise sales motion is the third area.

“We’re going to invest in our sales team. The goal is to increase that coverage by recruiting amazing salespeople who know how to work with the most sophisticated enterprises.” — Gabriel-James Safar, CEO, Tsuga

Q: What types of organizations is Tsuga specifically built for?

Tsuga is designed for enterprises that meet at least one of three criteria: sovereignty, meaning they need full ownership and control over where telemetry data is stored and what AI processes it; governance, meaning they need to enforce consistent observability practice and data quality across a large, distributed engineering organization; or scale, meaning they are generating telemetry volumes that have pushed them to the limits of what existing platforms can handle within budget. Safar notes that organizations meeting the scale criterion are often stuck on inadequate solutions simply because nothing else has been viable for them technically or financially.

“Our goal when we made Tsuga was to create a product that works for a very specific set of organizations. Sovereignty, governance, scale.” — Gabriel-James Safar, CEO, Tsuga

Resources & Documentation

  • Tsuga, bring-your-own-cloud observability platform built for enterprise sovereignty, governance, and scale
  • OpenTelemetry, vendor-neutral open-source standard for telemetry collection, explicitly recommended by Safar as the replacement for closed-source collectors
  • Databricks, data platform cited as a compatible downstream destination for telemetry data stored in open formats
  • BigQuery, Google Cloud analytics platform cited as a compatible downstream destination
  • Amazon Athena, serverless query service cited as a compatible downstream destination for open-format telemetry data

***

👇 Click to Read Full Raw Transcript

Swapnil Bhartiya: Today, telemetry volumes are exploding, other costs are spiraling out of control and traditional platforms can’t handle the rise of AI agents. By the time organizations realize they have lost control of the data, it’s already too late. They are trapped. Now SUGA is stepping in to help enterprises regain data sovereignty and take back control over their AI infrastructure. And today I have with us Gabriel James Safar, CEO of Tsuga. First of all, Gabriel James, it’s great to have you on the show.

Gabriel-James Safar: Thanks for having me.

Swapnil Bhartiya: It’s my pleasure. First of all, I would love to know a bit about the company itself. The name, when the company was created and what market shifts made you as well as your investors to completely rethink enterprise observability. So let’s talk about the story of the company.

Gabriel-James Safar: So my co founder Sebastian and I, we had been working together for quite a few years. Suga is our third company together. First one was a total failure. The second one got acquired by Datadog and we were in charge of a pretty large suite of products over there. And after our times had come, we left. What was clear from talking with large enterprises was that the Datadog model was great in a variety of ways. It was the best product for many customers. But at the same time it was not following a variety of shifts that we were seeing with notably larger organizations, organizations that have very large systems. And so that’s what gave us the idea of creating Tsuga. Tsuga as a world is a family of pine trees that grow in Japan and in the region of Seattle in the US and these are trees that have good properties for building things. And that’s why where the name came from. My wife is Japanese, so that’s why a Japanese name was nice.

Swapnil Bhartiya: Excellent. Thank you. The history and story of the company. Now the fact is that we are seeing a massive search in AI agents and autonomous system. And of course when it comes to observability, a lot of actually when I go to cnc, observability is I think one of the topics hottest topic these days of course open telemetry. The maturity and IT is becoming very critical piece of technology in modern world. Can you talk about how is AI fundamentally changing what enterprises actually need from their observed tools and what the current breed of tools is failing to provide them given the AI workloads.

Gabriel-James Safar: So if you put yourself in the shoes of any company with a large IT system, the first issue is AIs. So LLMs and agents are a new sort of service. So if you want to understand how your services are working in production, you need to understand these services because if you don’t, you won’t be able to analyze that. They are not going fast enough, they are not performing as expected, they are too expensive. So traditional observability problems. So problem number one, problem number two, if you want to benefit from AI, you will want your engineers to deploy a lot more often, right? A lot more code made by AI. So a lot more code turned onto production. But so if you do that, and if you want to identify that some of these new deployments are introducing issues, you cannot keep the same sample rate as what you used to. You need to keep a lot more data in order to see that. But IT systems kept growing in the past 20 years and so the telemetry volumes have kept increasing. So you have new services, you have an increase in the data volume. And last, in LLMs you have a problem that is just like in logs and in traces in the past. You have a lot of information that you probably don’t want to skip to let into the observability and not have the upper end upon, right? If you’re a company with healthcare in one of your customer is telling personal information about their health, that information will be in the telemetry of your LLMs. And so is it okay to have that information sent into a third party? Probably not. So these are the structural issues, increase of volume and increase in sensitivity and new sort of services that enterprise have to tackle. And on the other side, the opportunity is what AISRE can bring to the table. So you can see that as different topics that enterprise can seize in the age of AIs and how they need to deal with.

Swapnil Bhartiya: Can you also talk about if we just forget about AI for a second because AI is putting a different kind of strain on the workloads in observability in general. If you look at the whole observability space, if you look at open source side of it, you know, open sensors, open telemetry, you know, open tracing and now open telemetry has kind of become a big project in the space. Are you once again, if we ignore AI, what do you feel about the whole evolution of observability? And then if you bring AI into the picture, did AI accelerate the evolution of observability or that evolution that was already due? AI just kind of stepped in to make things faster. What I’m trying to understand is the evolution of observability as it was happening. Did AI change it, speed it up or made people rethink it.

Gabriel-James Safar: Indeed. What you’re evoking is the notion of data collectors, right? So the agents in the ancient meaning of the world, right? So these, so the observability agents, so the things that collect the telemetry have been closed sourced for a very long time and it created issues because it created a vendor lock in for enterprises using observability. And that was pretty bad. I think getting rid of these locked closed source agents has been a top priority project for the past five years in many, many enterprises. What AI is bringing to that is that the transition is much simpler, it’s much easier using AI to do the transition. It doesn’t make it instantaneous, right, but it makes it easier. So AI is helping enterprise become less vendor locked when it comes to the data collection of the observability on the other side, telemetry is extremely valuable. You can do a lot of things, of course, you can troubleshoot, you can do analytics, you can identify optimizations in your systems, but it’s also business data, right? And so if you can feed telemetry to your AI, you can unlock a lot of value by bridging it notably with business information. So I think that a project for many enterprise has been also to own that data, which was not possible with the products in the past. That’s why we brought bring your own cloud as well, right? Because bring your own cloud allows customers to own their data. And so if they want to feed it to their BI tools, to their own AIs, they just can do that because the data is there.

Swapnil Bhartiya: And since you mentioned they do want to own data and which has been case. But the whole thing, when you move to the cloud, we do talk about data gravity there. Once the data is in there, egress cost can make it very, very hard for you to move. But now we are talking a lot about data sovereignty in Europe. A lot of laws are coming in. AI sovereignty is being talked about because of this whole geopolitical conflict going on. Countries are very, very kind of skeptical of trusting each other. So they do want to move in data at the same time. We can talk about privacy and all those things. Can you talk about what is driving this change in priority for data sovereignty? Whether it’s AI, whether it’s geopolitical or evolution of technologies, and what role is observability and SUGA playing in this space?

Gabriel-James Safar: So I know we talk a lot about Europe in that domain, but Europe is not alone, right? Canada, Brazil, Japan, India, Australia, countries in the gcc, like KSA or uae, all these countries have jurisdictions that make it more and more important to choose where your customer data is being stored. And why? Well, because these countries understand that this customer data is very valuable and giving it away is not a good idea. And the same way these countries see that, the companies see that as well. So for me, where there is a big shift is that the notion of sovereignty is not only country level. You can see that also as company level. I want to be sovereign in the sense that I want to own the pieces of my stack and I want to own without having to pay a tax, my data. Because if I don’t, then all of a sudden a vendor can just use my data to feed my competitors. And that’s something that I don’t want as a company. So the notion of sovereignty, I don’t think we should hear that only as at country level. We should also see that at company level, notably for companies that are important. Right? If you’re a big company, you are a sovereign entity. So in that regard, the approach we have in SUGA is bring your own cloud, meaning that we are compatible with all the open source data collectors. So you can keep your own collectors opentelemetry open or others on one side, the data is stored in the entire processing, the entire observability system is in your cloud. Meaning that the data doesn’t leave, meaning that you keep the control, the data is yours. You don’t pay a tax to use it, right? You don’t pay an infra tax to use it. And on top of it, even when we look at AI, we work with the bring your own agent approach. So you can choose your AI systems and run it on top of the telemetry. And we provide a variety of products and systems to make your AI model that you approved. Because it’s very much a political decision. What AI and what’s the AI posture of a company? And you can just apply that to our system and we’ll harness it so that it works.

Swapnil Bhartiya: Let’s talk about cost a bit. A lot of organizations are kind of drowning in rising telemetry volumes and cost. Can you talk about why our traditional approaches to observability are not sustainable in this new AI era and how SUGA is also focusing on the cost aspect, if you will.

Gabriel-James Safar: So that’s what we discussed at the beginning, right? You need, you have more services, it keeps growing and you need to sample a growing amount of the data. So the volume of telemetry keeps increasing. At the same time, there is one thing that is fixed. And what’s fixed is not the amount of telemetry. What’s fixed is the budget. So if you’re an enterprise, you have a budget and you don’t want to blow it up. So for these reasons, historically, the solution was to tell you, okay, what about sampling even more? Even more? Always more. Always more as well, right? You sample and you keep in the end 1% of the data or 01% of the data. But then it defeats the purpose of telemetry because you never know when you’re going to need a metric where you’re going to need a trace. So the more you sample, the more you create operational problems for your teams because they need to spend a lot of time in reducing the volumes and the more likely it is that you won’t have the right data when you need it to do analysis or troubleshooting. So that’s why we wanted to flip the paradigm, right? If we were not bring your own cloud, we would need to sell you observability. It would have a cost on our infra and then we would need to sell it to you 5x more to have 80% gross margin. The goal is to flip that and to say we are going to make it work so that even if you have a huge amount of data, we’ll make it work so that the pricing allows you to keep maybe 10 times more data than what you could do with another system. And because it’s in your infra, potentially you can even negotiate better prices with your cloud vendor, which can have a good impact on your global cloud bill. So that’s the approach. The approach is instead of putting an infra tax, we want to design a system where you can have as much data as you need for your teams to go at the fastest speed possible.

Swapnil Bhartiya: Now when we talk about of course telemetry, of course the first word that or first term that comes in every open telemetry and of course open source. Looking at this geopolitical crisis, open source kind of become the universal language. It removes a lot of barriers to entry. But the beauty of open source is that it’s community maintained. It’s not controlled by a single vendor. That means you are not locked or you are on the mercy of that vendor. The problem is that open source can solve day one problem very easily. You can download the code, you can get it installed, but then day two becomes a big challenge. That’s why you need enterprise prayer. That’s why commercialization in open source is very, very important for the success of open source. Sometime you cannot have a puritan word. You may want everything to be open source, but you may have to have a mix of open source and proprietary. What is suga’s approach towards open source and observability?

Gabriel-James Safar: So we are strong believers in the fact that the data collectors on one side should be open source. Because these data collectors, these agents, you put them in your system, in your code. So if ever you want to leave your vendor, you need to be able to. So if it’s closed source, it’s creating the wrong pressure on the value at the other side. The storage format needs to be open source so that the data is not locked into our own ecosystem, but it can be used across your databricks, your bigquery, your Athena, etc. So we are big believers in the data should be open source through and through. And in the middle the goal is for us to make a very opinionated product. So on that front it’s harder to be open source and extremely opinionated. So for now at least the central piece of the product is not open source. But we ensure that there is a pressure for us to deliver because if we don’t deliver enough value, our customers have open source data collectors at the entrance, open source data at the exit, and so they can get rid of us and keep their data flow end to end even without us.

Swapnil Bhartiya: And that is the right approach to not forcing people to get logged in now, which is also a very good segue that as organizations are trying to transition away from legacy platforms, what is the biggest roadblock that they usually had that when you talk to them they talk about all those challenges problem and how do you folks help them get past that roadblock?

Gabriel-James Safar: The answer won’t surprise you. The biggest blocker is usually the people. You have people. Observability is used by a huge community within an organization. So many workflows depend on that. So usually the biggest blocker are the people. How can we empower the people then? Of course, so that’s the biggest broker, that’s the biggest thing that needs to be identified with the central team so that we can empower them and show them how with the transition, their life is going to be easier. In Suga, we have a lot of rules to help on governance of telemetry, govern the quality of data, the quality of observability, assets, enablement. So that’s how we can help the central teams do that. But that’s one piece. The second piece that is complicated is the data collection. Notably for organizations that come from products with closed source solutions. So on that we have forward deployed engineers who can help them change their collectors. So install OpenTelemetry agents to make it efficiently design the architecture. To give you an example, we support multi cluster. So if you are an organization, you want to keep a part of your data in the us, a part of your data in Europe, a part of your data in Brazil, you can do that and have a single interface in our product. What should be the data flow, what should go where? Is something very important where forward deployed engineers can help identify the bottlenecks, help design the architecture and implement it if that’s what’s going to make the organization more efficient. We like to say that we are not a SaaS, but we are a SaaS. So we are a SaaS in the sense that we are a software and a service. So we are forward deployed engineers. We have great partnerships with partners in the ecosystem and we empower them to bring a lot of value. So we focus on the human factor and of course a set of tools to make that more efficient so that transitions can be a success.

Swapnil Bhartiya: Now let’s talk about the growth of Suga. You folks raised I think 30 million if I’m not wrong. Talk a bit about with this new funding round, what is going to be the primary focus for investment and company growth, Engineering, product team, sales.

Gabriel-James Safar: So three investments. Number one is we need to observability. Even when you know what you need to do. Because we’ve been in the space for so long, my co founder, my head of product, many of our engineers that we have a good idea of things that we want to build. But even if you have a good idea, the world is changing and there are a lot of things to do. So we’re going to keep investing on the product. There is a lot of things we want to do and that we have not done yet. So that’s number one. Number two is we’re going to invest in our sales team. Currently we have salespeople in France, Germany, uae, us, uk. The goal is to increase that coverage by recruiting amazing salespeople who know how to work with the most sophisticated enterprises and to recruit the forward deployed engineers who can work with them. And last, we are going to invest in marketing in order to support this motion. That’s very classical, right? For a series A product, sales and marketing to support the sales.

Swapnil Bhartiya: When it comes to observability space, you mentioned a lot of names. It’s a very busy space, you know that it’s a crowded, busy space. Why should an organization look at SUGA when they want to solve their observability problems?

Gabriel-James Safar: That’s a very good question. So our goal when we made SUGA was to create a product that works for a very specific set of organizations and that’s the organizations who care about at least one of the following three sovereignty. Ensuring that they own their data through and through, that they control where the data is stored, where they control the AI that work on telemetry. So sovereignty is number one. Number two is governance. So ensure that observability is practiced correctly, company wide by empowering the central teams to make observability a shared language across all the organizations. And we do that through the product and through our enablement. And last for scale, we are best used with organizations that have a very large scale in volume of data, notably and who very often are blocked onto solutions that are not great but because it’s the only one that work for them in terms of organization or budget. And that’s where we can help. Right? So sovereignty, governance, scale.

Swapnil Bhartiya: Gabriel James, thank you so much for joining me today and sharing your insights. A lot of things are happening in the space, so I would love to have you back on the show. But I really, really appreciate your time today. Thank you.

Gabriel-James Safar: Thank you very much.

Swapnil Bhartiya: And for anyone watching who’s struggling to manage their telemetry data and cost, definitely check SUGA out and what this team is building and I look forward to chat with you folks again. Thank you.

How Finance Leaders Should Define AI Success | Peter Maloney, Azul | TFiR

Previous article