Traditional observability platforms detect failures and page an engineer. They do not fix anything. The engineer wakes up, consults the runbook, and applies the same remediation that was applied the last time the same error fired. Meanwhile, the platform has already exported your production logs to a third-party cloud, adding egress cost and expanding your data exposure surface. Research cited by Grafana puts observability spend at roughly 17% of total modern infrastructure cost, and the majority of that spend does not reduce mean time to resolution.
In this interview on TFiR, Ishay Yaari, Co-Founder and CEO at DataAgent, breaks down why the alert-and-wait model is architecturally broken, how DataAgent deploys an AI-native SRE inside a customer’s own cluster to remediate errors autonomously, and where the guardrails sit to keep autonomous action safe in production.
Guest: Ishay Yaari, Co-Founder and CEO at DataAgent
Show: TFiR
Here is what every SRE, platform engineer, and infrastructure architect needs to know.
Technical Deep Dive
Q: Why do traditional observability tools fail to fix production errors?
Ishay Yaari, Co-Founder and CEO at DataAgent, argues that observability is a feature, not a solution. Every major observability vendor has responded to the AI era by placing an agent on top of the same logs they already export from customer clusters to their own SaaS, then surfacing a dashboard. That approach does not reduce MTTR; it adds a token and licensing cost on top of a data egress bill while the error sits unresolved until a human intervenes. The fundamental problem is that the alert-and-wait model was never designed to act.
“Observability is basically a feature, it’s not a solution in today’s world.”
Ishay Yaari, Co-Founder and CEO, DataAgent
Q: What was the founding insight that led to DataAgent?
Yaari and Co-founder Nati Shalom, who previously built and sold Cloudify to Dell, identified that nearly every observability vendor was simply wrapping an AI assistant around exported log data. Their prior experience at Cloudify taught them to build topology maps of entire application ecosystems, from the application layer down to infrastructure services. Applying that topology approach to error detection inside the cluster, without exporting logs, was the architectural starting point for DataAgent.
“We realized that almost 100% of the observability tools, what they do in order to serve their customers in the AI area, is just putting an agent on the same logs that they keep exporting from the clusters to their site.”
Ishay Yaari, Co-Founder and CEO, DataAgent
Q: How does DataAgent detect errors without shipping logs outside the cluster?
DataAgent listens to more than 20 signals inside the cluster, including Google Golden Signals metrics, and maps them against a continuously maintained topology view of the application ecosystem. Because the detection layer runs inside the customer’s own environment, no production data is exported to a third-party SaaS to trigger an alert. Yaari states the platform can predict errors at 100% accuracy using this signal-plus-topology approach, and that signal injection into the topology view makes it possible to understand exactly what changed, not just that something changed.
“Instead of exporting log outside the cluster, we listen to more than 20 different signals. We can predict in 100% accuracy there is an error.”
Ishay Yaari, Co-Founder and CEO, DataAgent
Q: Why does alert fatigue get worse when observability tools export more logs?
Yaari explains that many exported logs report the same error repeatedly, so the volume of alerts a team receives does not correspond to the number of distinct problems. DataAgent addresses this by regrouping redundant signals inside the cluster before any data leaves, collapsing noise at the source rather than at the dashboard layer. The result is a smaller, higher-signal alert queue that reflects actual distinct incidents.
“Many of the logs that are being exported outside of the cluster report the same error. We found a way to regroup them, to eliminate a lot of noise when it comes to alert fatigue.”
Ishay Yaari, Co-Founder and CEO, DataAgent
Q: What is the real cost of shipping observability data outside the cluster?
Yaari references Grafana’s own research, which puts observability cost at approximately 17% of total modern infrastructure spend. DataAgent’s internal research indicates that organizations can reduce that bill by up to 90% by stopping log exports and relying on in-cluster signal detection instead. The argument is that cloud providers such as AWS already surface the data needed to detect errors natively; paying a second vendor to re-ingest the same data is a redundant tax.
“About 17% of the cost today is related to observability when it comes to the modern infrastructure. That’s a bill that you can reduce by up to 90% if you just stop shipping logs outside the clusters.”
Ishay Yaari, Co-Founder and CEO, DataAgent
Q: How does DataAgent operate as an AI-native SRE inside a company’s own infrastructure?
When a signal detects an error, DataAgent alerts the SRE, performs root-cause analysis, and proposes a remediation plan. The user can configure specific error types to require a dry run before the fix is applied, or to run in full autonomous mode. Over time, errors the SRE has confirmed as safe are added to a fault catalog, and DataAgent applies those fixes automatically without human intervention going forward.
“As the user learns how to trust the remediation plan that we offer, every time the user decides to dry run and fix and let us fix the error, he can also mark that error to be a candidate for full autonomy.”
Ishay Yaari, Co-Founder and CEO, DataAgent
Q: What percentage of production errors can be fixed without human intervention?
Yaari cites internal research showing that more than 80% of errors in modern infrastructure can be resolved with a restart or a rollback. DataAgent applies this as a circuit-breaker logic: remediate first, then investigate, analogous to taking an over-the-counter remedy before deciding whether to go to the emergency room. That 80% figure is the foundation of the platform’s claim that most alert volume is noise that can be eliminated before it reaches the SRE queue.
“More than 80% of the errors in today’s infrastructure can be simply a restart or rollback. You apply the fix, then you start investigation after you apply this circuit breaker.”
Ishay Yaari, Co-Founder and CEO, DataAgent
Q: What is a fault catalog and how does it enable autonomous remediation?
The fault catalog is a maintained record of every error type the SRE has confirmed DataAgent can fix autonomously. Each time an engineer approves a dry-run fix, that error pattern becomes a candidate for full autonomy and is added to the catalog. The catalog compounds over time: as more error types are confirmed, the proportion of incidents handled without human intervention grows.
“This fault catalog is basically allowing us to prevent errors, to fix errors in post production. And this is where the story gets really interesting.”
Ishay Yaari, Co-Founder and CEO, DataAgent
Q: What is the DataAgent digital immune system and how does it connect post-production and pre-production?
Once the fault catalog accumulates enough post-production remediation patterns, DataAgent can use those patterns to flag pre-production code that is likely to introduce known error types before it ships. Yaari frames this as a cure-and-vaccination cycle: the platform provides automated remediation in production and proactive prevention signals to the engineering team during development. The two loops reinforce each other as the catalog grows.
“If we learn how to solve errors in post production, we can also alert the DevOps team that the code about to ship from engineering might cause errors. We call it a digital immune system: a cure for post production and a vaccination for pre production.”
Ishay Yaari, Co-Founder and CEO, DataAgent
Q: Can DataAgent run alongside Datadog or Dynatrace without replacing them?
Yaari explicitly states that no large enterprise will rip and replace Datadog or Dynatrace on day one, and DataAgent is designed to overlay on top of existing tools. In this configuration, DataAgent handles the 80% of errors that can be auto-remediated in-cluster, preventing those logs from ever reaching the downstream observability platform. Datadog or Dynatrace then receives only the 20% of incidents that genuinely require deeper investigation, which also reduces the ingestion bill for those platforms.
“We can basically prevent the log that goes outside the cluster into Datadog by 80%, just by overlaying Data Agent on top of Datadog.”
Ishay Yaari, Co-Founder and CEO, DataAgent
Q: What measurable ROI does the remediation-first approach produce beyond faster MTTR?
Beyond MTTR reduction, DataAgent runs a discovery mode that maps all logs currently open across the customer’s stack, including overlapping logs from multiple tools, and recommends which ones to shut down with an explanation of the consequences. Yaari states that accepting these recommendations can reduce total observability spend by up to 90% from current levels. The topology-level view makes it possible to identify redundancy that is invisible when looking at individual tool dashboards in isolation.
“We also give some context. We explain what is the meaning of shutting down these logs. That saving can go up to 90% from what they currently pay today to these tools.”
Ishay Yaari, Co-Founder and CEO, DataAgent
Q: What are the AI guardrails that keep autonomous remediation safe in production?
DataAgent uses a two-layer engine architecture. A deterministic layer maintains the topology map of all services, tools, logs, and infrastructure layers, with no AI inference involved in that mapping. An adaptive LLM layer handles root-cause analysis and remediation proposal generation, where some degree of inference is necessary. Guardrails on the LLM constrain it strictly to infrastructure self-healing tasks; Yaari describes a chat interface called Otto that will answer infrastructure questions in detail but refuses entirely out-of-scope queries.
“The secret sauce is the balance between the deterministic side of the technology and the adaptive side of the technology. If you ask our chatbot how do you make coffee, it will answer you, that’s not my business.”
Ishay Yaari, Co-Founder and CEO, DataAgent
Q: Does DataAgent use its own LLM or can enterprises bring their own?
Yaari states that DataAgent offers both options. The platform has its own LLM, and customers can also supply their own LLM if they have data residency, sovereignty, or vendor preference requirements. This bring-your-own-LLM capability is positioned as a direct response to enterprise security concerns about AI systems accessing production infrastructure data.
“We use our own LLM or we offer our customers to bring their own LLM.”
Ishay Yaari, Co-Founder and CEO, DataAgent
Q: How do enterprise security and compliance concerns affect autonomous SRE adoption?
Yaari reports that even access to a development environment now triggers security review, with prospective customers asking for SOC 2 certification before permitting any integration. The concern is not irrational: AI systems with write access to production infrastructure represent a different risk category than read-only observability tools. DataAgent’s response is the dry-run and fault-catalog model, which requires explicit human approval before any error type is elevated to autonomous remediation.
“In our previous cycles the answer was, hey, let’s try it today. Now they want to understand how your system works. Do you have SOC 2 before you get into my dev environment?”
Ishay Yaari, Co-Founder and CEO, DataAgent
Q: What does autonomous SRE look like in practice for a team using DataAgent today?
Customers see a dashboard showing how many errors were resolved automatically in the past 24 hours, how long each resolution took, and which errors are pending human review. SREs use the Otto chat interface to query what happened in a given time window and are presented with a prioritized queue of incidents that require their judgment. Yaari states that customers report saving hours of SRE time per day by focusing human attention exclusively on the errors the system cannot yet handle autonomously.
“The real value is the separation between the errors that are self-driving and save hours every day on average for every SRE, and allowing them to focus only on the errors that require their intervention.”
Ishay Yaari, Co-Founder and CEO, DataAgent
Q: How do engineering teams decide which fixes to let autonomous agents handle versus requiring human sign-off?
Yaari is direct: there is not yet enough industry data to give a reliable benchmark. The autonomous SRE category is at most 12 to 16 months old for most entrants, and every customer’s infrastructure and risk tolerance is different. The practical answer is that teams start by approving dry-run fixes manually and gradually expand the fault catalog as trust accumulates. Yaari expects meaningful benchmark data to emerge within the next year as more teams operate these systems at scale.
“Every customer is different, every infrastructure is different. The most important message is that we allow full flexibility. Slowly and surely you add more and more errors to the fault catalog and allow us to automatically fix them for you.”
Ishay Yaari, Co-Founder and CEO, DataAgent
Q: Where does autonomous SRE stand today compared to AI-assisted observability?
Yaari draws a sharp line between AI-assisted observability, which pulls logs out of clusters and returns natural-language summaries, and autonomous remediation, which acts inside the cluster to resolve errors without human intervention. He characterizes the former as the dominant model for the past three years and the latter as a category that is only beginning to emerge in the past six months. The transition is being driven, in his view, by the recognition that assistant AI increases token and licensing cost without reducing the error backlog.
“In the past six months you start seeing more and more technologies start thinking that maybe assistant AI is not the answer because it’s not going to fix the error for you, it’s just going to increase the cost.”
Ishay Yaari, Co-Founder and CEO, DataAgent
Resources & Documentation
- DataAgent, AI-native autonomous SRE platform for in-cluster error detection and remediation
- AWS CloudWatch, native AWS monitoring referenced as sufficient signal source for error detection without third-party log export
- Google Golden Signals, the four monitoring signals (latency, traffic, errors, saturation) used by DataAgent’s deterministic detection layer
***
👇 Click to Read Full Raw Transcript
Swapnil Bhartiya: When it comes to observability tools, they are good at exactly one thing. They tell you that something has broken, then they page a human and wait. The fault stays exactly where it was until an engineer wakes up, opens the runbook and retypes the same fix that they typed last time. And to get even that much visibility, most platforms need your production data shipped outside your own infrastructure first. That’s cost and obviously exposure is stacked on top of your workflow. That is already too slow. Data Agent is trying to flip that model. Instead of watching and waiting, it acts running as an AI native SRE inside a company’s own environment, resolving false error in real time with guardrails, keeping every fix safe and also keeping your data in your own environment. And joining me today is Ishay Yari, co founder and CEO of Data Agent to talk about this whole concept. First of all, Ishaya, it’s great to have you on the show.
Ishay Yaari: Thank you very much. Thank you for having me. It’s a pleasure being here.
Swapnil Bhartiya: It is my pleasure since you’re a co founder of the company. And then we look at this problem and this is actually becoming a very big problem. I just came back from Splunk and the focus is more about don’t bring your data agents because that is, you know, cost but also privacy, security and tons of stuff there. What I want to understand is that what gaps in traditional observability tools were there that push you to start Data Agent? Just talk about the story idea behind the company, a little bit about my background. You know, about three years ago we ended a very successful journey with Cloudify. I partnered with Nati Shalom in the past, this is our second journey and basically Cloudify got acquired by Dell. And then after the acquisition Nati, my partner today spent his three years at Dell and I changed side to the private equity ward and worked with a London based private equity called Alicorn. And throughout my journey at Alicorn I got exposed to many many companies that touched different angle of infrastructure from CICD all the way to the infrastruct including the CI CD to CodeFresh and Anodot and Glassbox who captures user sessions which is again that’s where I expose to the datadog world and Dynatrace and all these giant companies that keep recording sessions and logs and that’s their core business. So about a year ago I regrouped with Nati and we started brainstorming what else can we do? Can we partner again and bring something new to the market? And as an X Clarify and X infrastructure automation people. We identify another big gap in the industry with the new AI area. We realized that almost 100% of the observability tools, what they do in order to serve their customers in the AI area is just putting an agent on the same logs that they keep exporting from the clusters to their site. And we realized that it’s not only helping the customers, it’s just overloading them with more cost. Because an AI assistant reading logs and give you a nice dashboard is not going to fix the error, it’s just going to increase your bill at the end of the day. So we start brainstorming what can we do differently in order to really bring new efficiencies into the observability space? How can we ensure application runtime beyond observability? And then we figure out that observability is basically a feature, it’s not a solution in today world if you take an agent and if you deploy it inside the cluster. What we learn to do best at cloudify was to create a topology view of the entire application ecosystem from the application layer, from the code all the way down to the infrastructure, the different services that allow the application to run. Then we realized that what if we create a discovery, a topology view of the application ecosystem and then instead of exporting log outside the cluster, we listen to more than 20 different signals, you know, for example Google Golden Matrix, we can predict in 100% accuracy there is an error. Now if you inject the signal into the topology view, we realize that we can understand what has changed. Not only to alert the user there is a change or something happened and then we start thinking, okay, so if we have the ability to understand what is happening inside the clusters, and if we can predict in 100% accuracy the errors, and if we can prevent exporting log outside the system, there’s probably a way we can also apply remediation into most of the errors. Because we did a research and we find out that more than 80% of the errors in today infrastructure can be simply a restart or rollback. Can apply the fixed, then you can start investigation after you apply this circuit breaker. Very similar to human being. If you have a fever, you don’t go to the emergency room, you try an anvil, then if the problem continues, then you go to the er. That’s the same logic we apply when it comes to the application infrastructure. So we did some research, we start playing with the infrastructure and the results were amazing. We were able to reduce drastically the noise inside the cluster. And we find out that by maintaining the blueprint of the application in the SaaS level and when asking the agent to report only the signals that detect errors, we can basically save not only time, but also cost drastically because we don’t force our users to pay double tax for the same data that they generate from aws. For example, if they run on aws. AWS basically give you everything you need in order to detect the errors. And you don’t need to pay double penalty to the datadog or the Dynatrace or even Grafana cloud of the world. It’s enough to send a signal and allow the platform to do a research for you where exactly what’s the source of the error? So that was phase one of the realization that there is a different approach in today’s technology that allow us to not only to observe but to remediate. And that’s why we came up with remediation. First we first tried to remediate, then we investigate and then we give you tools to solve the errors.
Swapnil Bhartiya: Can you also talk a bit about the role the whole alert fatigue that engineers deal with cost of shipping of data outside their enterprises?
Ishay Yaari: Okay, so it all related. It’s all part of the same story obviously. So we found out that when we can prevent exporting log outside the clusters, we basically can provide our customers a better, a more secure solution. We don’t require you to teleport data outside your cluster. As far as the alert fatigue, we found out that today many of the errors, many of the logs that are being exported outside of the cluster report the same error. So when we figure out how to resolve this, we found a way to regroup them to eliminate a lot of noise when it comes to this alert fatigue. That was second part of the story. The third part, we also uncovered the true cost of exporting this logo. And the big surprise and even reported by Grafana, about 17% of the cost today is related to observability when it comes to the modern infrastructure. And that’s a lot. That’s basically a bill that you can reduce by up to 90% if you just stop shipping logs outside the clusters.
Swapnil Bhartiya: Now let’s also discuss how does Data Agent actually operate as an AI native SRE inside a company’s own infrastructure? And if we can also talk about the guardrails, how do you also make sure that autonomous fixer stays safe in production?
Ishay Yaari: So by default when there’s an error inside a cluster and one of the 20 different signals detects an error. We alert the operator, the SRE that there was an error. Now by default we alert and we propose, we do a root cause analysis. We propose a remediation plan by default. And then the user can basically configure data agent to decide what errors they can let us self remediate or what they need to dry run before they apply the fix. So in other words, as the user learns how to trust the remediation plan that we offer in related to the errors, every time the user decides to dry run and fix and let us fix the error, he can also mark that error to be candidate to a full autonomy. And that basically what we do behind the scene, we maintain a catalog of all the errors that the user basically has confirmed that data agent can fix for them. And think about this fault catalog is basically allow us to prevent error to fix errors, sorry, to fix error in post production. And this is where the story gets really really interesting is because once we learn how to fix errors in post production and the SRE learn how to trust these remediation plans, this fault catalog now serving post production can also be feeding pre production. Think about it, if we learn how to solve errors in post production, we can also alert the DevOps team that the code that is about to ship from engineering might cause errors. Because we learn about the nature of these errors in post production. And if you think about it, this is a cycle that feeds itself. We learn how to fix error in post production, the user tends to trust us and let us self drive these resolutions. But we also propose prevention solution for the engineering and this cycle we call it a digital immune system. We provide our customers a cure for post production and a vaccination for pre production. To answer your question in a short way, we allow the user to configure to run a data agent even side by side next to their existing observability. Because we understand there’s no large enterprise that is going to rip and replace Datadog or Dynatrace in day one. So we developed the product in that way that can also alert the user when the error justifies datadog to do the root cause analysis and allow us just to filter the noise. And if we go back to the second question you asked me, after we learned that 80% of the errors are basically just noise that you can roll back or restart and fix them, we can basically prevent the log that goes outside the cluster into datadog by 80% just by overlaying Data Agent on top of Datadog. So again the message is that we are very flexible as far as how the user can use Data Agent in production.
Swapnil Bhartiya: Beyond just fixing failures faster. What other measurable results are you seeing in MTTR cost saving and developer productivity with this remediation first approach?
Ishay Yaari: That’s a good question. So another angle that is a very interesting angle and I touched a nerve in the previous question, but now I’ll give you the full answer because we understand the infrastructure from actually not the infrastructure, the application ecosystem from the application layer all the way to the infrastructure. This topology view also allows to detect the different logs that are open by default by the other observability or other tools that the customer is using. So part of the solution or part of the benefit or ROI that we can propose to our customers is to run this discovery mode and basically give them a map of all the logs that are currently open by default and are not necessary. We’re not just flagging these open logs that are open. We also give some context. We basically explain the user what is the meaning of shutting down these logs. So the ROI here is that once the user accepts our recommendation, they can drastically lower their cost by just shutting down logs that are overlapping with each other. And based on our research, that saving can go up to 90% from what they currently pay today to these tools.
Swapnil Bhartiya: Now to ask that how do you see the whole autonomous agent ecosystem will evolve is a question that even I mean that’s what we learned. So I’m not going to ask that. But what I do want to ask is that looking at first of all the autonomy of AI agents, they are doing a lot of things on their own. They are making critical decisions. Sometime human is in loop. Sometimes humans is not in loop. At the same time we are also hearing from these big three players, the whole hugging face, you know, incident that happened that we need to slow AI down. I don’t think it’s a good idea to slow something down. The better idea is to build better guardrails and stuff like that. What kind of concern do you see from organizations that, you know what, I am not comfortable plugging AI into my data, into my system. Apps are, they can come and go. Data is very, very critical. And then what kind of confidence you give them that they are willing to take that risk?
Ishay Yaari: That’s a good question. And we see this a lot in the conversation we are having with the first customers we onboard. They all have concerns. Even in the past when you had a new tool and you go into a customer and you ask them, hey, take it, try it in your dev environment. In our previous cycles the answer was like, hey, let’s try it today. They want to understand how your system works. Do you have SOC 2 before you get into my dev environment, let’s make sure that everything is protected. And I think this fear is valid. To answer your question, I think it’s a mix, it’s a balance. If I talk about how we do it. So the way I frame it is that we kind of build a private case of Claude specifically for infrastructure self healing. It means that we have a very deterministic engine that understands all the different services and tools and logs and layers that you currently have. It’s not, there’s no guesswork here, there’s no AI work here. It’s a very deterministic engine that understands. But we also have an adaptive side of our engine that once the error happens and we start root cause analysis and we need to propose a remediation, that remediation, we also understand that we need an adaptive engine that can do some tests in order to get to a resolution plan. And I think the secret sauce is the balance between the deterministic side of the technology and the adaptive side of the technology. And in addition to the fact that you need to put some guardrails on top of the LLM, you know, we use our own LLM or we offer our customers to bring their own LLM. But in a nutshell, if you log into Datadog today and ask our chatbox, Otto, that’s how we call our chatbox. If you ask our chatbox what happened in my stack in the last 24 hours, you will get a response from our side that looks like Grafana. We generate a Grafana output at no cost. We don’t need Grafana. Our engine can self create this mapping for you. But if you ask Otto the same chatbox how do you make coffee? He will answer you, that’s not my business. So I’m trying to add some humor here, but I think it’s all about how you control the LLM and how you put some guardrails on what it is allowed to do and what it cannot do for you. I think this is the balance that we all aim for.
Swapnil Bhartiya: Is it possible for you to walk us through a real example of a fault data agent resolved on its own, all the way from detection to fix. You can name the customer or it was internal, just to give our viewers an idea how it actually works and if there are some use cases where it worked.
Ishay Yaari: I’m trying to think about an answer that will satisfy both the technical side and the high level side of the business. But in a nutshell, when there is an error, without getting into the technicality of the error or the layer of the error, every customer on average has about between 50 to 1,000 errors a day for many many reasons. So the use case that we always talk about when we go and speak to other customers is that we show our active customers. They have a dashboard and they have a view of exactly how many errors were resolved automatically in the past 24 hours, how much time it took for each error to be resolved, what is the pending errors that is requiring a human in the loop. So when an SRE is logging or basically using their own interface to ask Data Agent what happened in the past hour, then they have a queue that is requiring their intervention. And just by doing that we are learning that we are saving our customers hours every day just by allowing them to focus only on the problems that require their human intervention while others are basically running on self mode. And I think this is the example that comes to my mind here. The real value is the separation between the errors that are self-driving and save hours every day on average for every SRE and allowing them to focus only on the errors that require their intervention.
Swapnil Bhartiya: And what percentage. How do enterprises decide which fix they are comfortable letting agents do versus where humans need to sign off? How do they decide how much control they have and what you have seen how much percentage they let autonomous agent do it versus they’re still like, hey, we need human in loop.
Ishay Yaari: So to really give you an honest answer, I think we will need to wait another six months. I think because the industry is too premature for all of us to generate benchmark. Even if you compare us to the other new entrants in the market, they exist only within 12 to 16 months. So I don’t think we have enough benchmark to tell you exactly what type of errors can be self automated versus the one that requires the human loop. But I can also tell you that every customer is different. Every infrastructure is different, every customer has different concerns. So I don’t see a common thread yet. And that’s why I think the most important message is that we allow full flexibility. We flag the incident, we start root cause analysis, we show you exactly what we understand the steps need to fix these errors and then you decide and slowly and surely you add more and more errors to the fault catalog and allow us to automatically fix them for you. And unfortunately this industry, the autonomous SRE industry is too new to tell you what the benchmark is when it comes to the error types and the average resolution or what can be self healing versus human in the loop. I think we need to regroup, I don’t know, within the next year and rediscuss this. It’s going to be a very interesting topic. I guess that will determine the future. I can tell you that autonomous SRE is not a new word. But what we can see is that from the day people start using autonomous SRE, I don’t know, three years ago, most of the work was mainly assistant AI. Again going back to the very first question you asked me, pulling log outside the clusters, put an agent on the logs and give you fast answer what happened. Not what has changed, just how many errors. And I think in the past six months you start seeing more and more technologies start thinking about maybe assistant AI is not the answer because it’s not going to fix the error for you, it’s just going to increase the cost because tokens and LLMs also cost. And I think the autonomous side of the equation is literally starting right now. So it’s too soon to talk about benchmark if I want to be responsible.
Swapnil Bhartiya: Ishai, thank you so much for sharing how Data Agent is moving observability from alerting to actual remediation. Thanks for sharing some examples here as well. And for those who are watching, please go and check Data Agent to understand how it can solve your problem as well. And back to you, thank you so much and I look forward to chatting with you again.
Ishay Yaari: Thank you for having me, it was a pleasure.





