AI Infrastructure

Gremlin Reliability Scores Now Live Inside Dynatrace: What SRE Teams Need to Know | TFiR

0

AI-generated and agentic code is reaching production with less human review than any previous software delivery model. Classic observability platforms tell you what already broke. They do not tell you what will break when latency spikes, a dependency fails, or an agentic pipeline deploys code that has never been stress-tested. The gap between what monitoring shows and what production actually does under pressure is where most P1 incidents are born.

In this interview on TFiR, Kolton Andrus, CEO and Founder at Gremlin, and Philippe Deblois, Global Vice President, Solutions Engineering at Dynatrace, cover the integration that brings Gremlin reliability scoring and resilience testing natively into the Dynatrace platform, and explain how engineering and leadership teams can use it to prevent failures before they impact customers.

Guest: Kolton Andrus, CEO and Founder at Gremlin
Guest: Philippe Deblois, Global Vice President, Solutions Engineering at Dynatrace
Show: TFiR

Here is what every platform engineer, SRE, and engineering leader needs to know.

Technical Deep Dive

Q: How is AI changing the reliability risks companies face in production?

Kolton Andrus, CEO and Founder at Gremlin, and Philippe Deblois, Global Vice President, Solutions Engineering at Dynatrace, both identify the same core problem: AI and agentic development pipelines are accelerating code delivery while reducing the depth of testing before production. Code is shipping faster, complexity is increasing, and the traditional questions of whether a service is running and performing must now be joined by new questions around accuracy and correctness specific to AI workloads. The honeymoon phase of AI adoption carries a real trade-off: vulnerabilities and defects that have not been caught are making it into turbulent production environments.

“We’re shipping code really fast. The complexity is increasing. But how can we make sure that that code is going to be reliable? With AI workloads now we have additional questions. Is it accurate? Is it working in the way that we expect it to?” — Kolton Andrus, CEO and Founder, Gremlin

Q: Can you trust observability data when AI is involved in generating or operating systems?

Deblois explains that Dynatrace addresses this through deterministic AI, not exclusively through LLMs. Deterministic AI produces answers that can be trusted and acted upon with confidence, forming the foundation layer of the platform. LLMs and probabilistic models are used on top of that foundation to help present data and smooth user interfaces, but they are not used to process raw observability data or perform deep analysis. This architectural separation is what makes the platform’s outputs reliable enough to act on.

“Deterministic AI is a key component of observability and a foundation for leveraging probabilistic models like LLMs.” — Philippe Deblois, Global Vice President, Solutions Engineering, Dynatrace

Q: What is the right approach to using AI in resilience and reliability tooling without introducing new risk?

Andrus describes Gremlin’s approach as a prompt-and-confirm model: LLMs are used to simplify how users interface with results, not to process the underlying data or make autonomous decisions. The core analysis runs on machine learning and deterministic methods built on millions of experiments and tens of thousands of systems analyzed over Gremlin’s lifetime. That historical dataset is distilled into credible, actionable recommendations. LLMs help surface those recommendations in plain language and assist in identifying the right APIs and actions, but no action is taken blindly based on a language model output.

“We’re going to use the LLMs for helping smooth out the interface and find the right actions to take. But we’re going to do it in a way where we prompt and we confirm. We don’t just blindly take action and believe what’s said.” — Kolton Andrus, CEO and Founder, Gremlin

Q: Why do resilience testing and observability need to be in the same platform?

Deblois frames classic observability as a retrospective tool: it tells you what has already happened. Gremlin enables a proactive model by letting teams create controlled failure scenarios and measure how production systems actually respond before an incident occurs. When both capabilities live in the same platform, teams can correlate experiment results with live telemetry, understand knock-on effects across dependencies and response times, and fix issues before they cause customer impact. Andrus notes that the cost of resolving issues early is substantially lower than managing them when they are affecting live production environments at scale, and AI in the stack compounds that cost if left unaddressed.

“Classic observability is only what has happened, not what could happen. Gremlin gives us the opportunity to go out and create these scenarios and see how the system actually responds to those.” — Philippe Deblois, Global Vice President, Solutions Engineering, Dynatrace

Q: What does the Gremlin app inside Dynatrace actually deliver, and how does it work?

The Gremlin team built a native application inside the Dynatrace platform using the Dynatrace app framework, which provides role-based access, permissions, security, and a UI that matches the rest of the Dynatrace experience. This app surfaces Gremlin reliability scores directly inside Dynatrace dashboards, so users do not need to context-switch between tools. It also exposes the output of Gremlin’s reliability intelligence, which is the machine learning-based analysis that runs when a test fails and tells the team why it failed and how to fix it. Triggers can be initiated from within Dynatrace to kick off Gremlin tests or scenarios, closing the loop between observation and experimentation in a single workflow.

“They can be in a dashboard and they can have reliability scores from Gremlin right from within their flow and continue to use the platform in the way they’re used to.” — Kolton Andrus, CEO and Founder, Gremlin

Q: How does the Gremlin and Dynatrace integration support leadership visibility and organizational accountability for reliability?

Deblois identifies leadership buy-in as a prerequisite for reliability at enterprise scale because reliability cannot be fixed by a single team. The Gremlin app inside Dynatrace enriches existing dashboards and reporting workflows that leadership already uses, making it straightforward for executives to see reliability scores, track risks, and hold teams accountable. The integration also solves a cultural incentive problem: teams that do proactive reliability work and prevent incidents are often invisible because there is no outage to measure. The data generated by Gremlin experiments gives those teams concrete evidence of what was prevented, enabling organizations to recognize and reward proactive engineering the same way they reward incident response.

“We often reward the folks that are great at firefighting and fixing the problems. But what about the team that does this work proactively and never ends up firefighting? We need proof. We need evidence.” — Philippe Deblois, Global Vice President, Solutions Engineering, Dynatrace

Q: What does day one look like for a Dynatrace customer installing the Gremlin app?

For existing Gremlin customers, setup begins in the Gremlin hub where the app is downloaded and the integration is configured. Andrus describes the process as simple, with the result being a shared language and unified UI between both platforms from day one. On the Gremlin side, significant engineering work has been done to make Dynatrace easy to connect: monitors and alerts can be pulled in quickly, tied directly to experiments, and used as both safety nets and feedback mechanisms during test runs. Deblois emphasizes that the design goal throughout was minimum time to value, so teams can begin seeing results without extended onboarding cycles.

“We want to make it very seamless for people that are already customers of both to be able to quickly get value, quickly set it up, quickly integrate, help them identify which monitors and alerts are the right ones to use.” — Kolton Andrus, CEO and Founder, Gremlin

Q: Can you walk through a real example of the integration catching a reliability risk before it caused an outage?

Andrus describes a large insurance company that used the Gremlin and Dynatrace integration during a cloud migration to test scenarios such as simulated outages and high latency conditions. The testing uncovered risks and potential P1 incidents before the migration went live. The migration completed with zero outages, which the company presented as a major success story at Dynatrace’s annual user conference. The engineer who led the effort was subsequently promoted. Deblois also describes the standard operational pattern: a team member arrives Monday morning, sees a new risk flag on a production service in their Dynatrace dashboard, drills into the Gremlin data for root cause details, applies the fix, reruns the experiment, confirms it passes, and closes the loop before any customer is impacted.

“This migration ended up with zero outage. It was one of their big success stories. And the lady that spoke was actually promoted as a result of this.” — Kolton Andrus, CEO and Founder, Gremlin

Q: How do Gremlin and Dynatrace handle support when something goes wrong with the integration?

Dynatrace customers can raise issues directly with the Dynatrace team, which coordinates closely with Gremlin across all levels of the organization including R&D. Andrus describes a shared Slack channel between the two teams as the practical escalation path for real-time issue resolution. Both teams have been committed to ensuring a smooth customer experience, and the partnership extends to active collaboration on how apps should be built on the Dynatrace platform and how integrations can be continuously improved based on feedback from the field.

“We got the secret bad signal line set up, the shared Slack channel. If something goes wrong, we’re talking about it and we’re fixing it.” — Kolton Andrus, CEO and Founder, Gremlin

Q: Where is the Gremlin and Dynatrace partnership heading as AI adoption accelerates?

Andrus describes the next frontier as ensuring that agentic systems, which are increasingly writing and deploying code autonomously, are subject to the same reliability feedback loops that human-driven CICD pipelines require. The goal is to feed reliability analysis and experiment results back into agentic pipelines so that code arriving in production has already been vetted: alerts and monitors are configured correctly, and the system has been tested against the types of conditions it will encounter. Deblois frames this as wanting innovation speed and reliability simultaneously, arguing that creating safety nets and feedback loops inside agentic workflows is the mechanism that makes both possible without forcing a trade-off.

“We want to have people moving quickly and taking advantage of AI, and we want them to do that without having a lot of outages, without having a lot of defects.” — Kolton Andrus, CEO and Founder, Gremlin

Q: What guardrails should teams put in place as AI adoption accelerates, and what are the risks of over-relying on LLMs?

Deblois and Andrus both return to the same principle: not all AI should be LLMs or generative AI, and teams that treat every problem as an LLM problem are introducing unnecessary risk. Machine learning and deterministic AI techniques developed over decades remain highly valuable for providing accurate, auditable answers in observability and reliability contexts. As agentic systems become more autonomous, Andrus identifies explainability and auditability as the critical next evolution of observability: teams need to be able to prove that an autonomous system did the right thing, not just observe that it appeared to. Deblois adds that explainability is directly tied to the long-term durability and trustworthiness of the systems being built.

“Having that visibility and having that explainability and auditability is really the next evolution of observability.” — Philippe Deblois, Global Vice President, Solutions Engineering, Dynatrace

Resources & Documentation

  • Gremlin, chaos engineering and resilience testing platform with reliability scoring and reliability intelligence
  • Dynatrace, observability and AI-powered analytics platform with native app framework for partner integrations

***

👇 Click to Read Full Raw Transcript

Swapnil Bhartiya: Now, when it comes to AI, the general perception is that adding AI, adding AI tools, AI agents, make your system smarter. It’s true, it does make it smarter, but it also makes them more fragile, more AI only deepens the blind spots. And we all know AI hallucinate. It can even lie. So you don’t even know what is going on. And of course, when it comes to observability, you do need to see what is breaking and prove that your systems can survive it. But AI kind of blurs the line. Now, Gremlin and Dynatrace are partnering to bring resilience testing and reliability scoring into one dashboard to make life easier for SysAdmins DevOps teams. And today we have with us two guests once again, Colton Andrus, CEO and founder of Gremlin, and Philippe Deploy, VP of Solutions Engineering at Dynatrace. Philippe Colton, it’s great to have you both on the show.

Kolton Andrus: Thank you so much. Glad to be here.

Philippe Deblois: Always a pleasure.

Swapnil Bhartiya: It’s my pleasure, actually. Of course we are going to talk about what you folks are doing, but before that, can we also talk about how is AI changing the way companies need to think about resilience? And at one end you can see all these tools can make you more resilient, but at the same time, it can also be part of the problem.

Philippe Deblois: I think what we’re seeing is, you know, anytime we go through an innovation cycle, there’s a lot of inflated expectations, there’s a lot of great gains to be had, and we’re kind of in that honeymoon phase where everything seems amazing, everything’s moving quickly, we’re able to just start having bots write code for us, we’re able to fully automate our deployment pipelines, we’re able to do this world of agentic development where they’re able to get things done quickly and out the door. But what is the trade off there? What are the side effects? And I think one of those are, is that a lot of that code that’s making it out the door hasn’t been tested as fully. And there’s likely to be vulnerabilities, there’s likely to be Mrs. In that code making it out the door that need to be tested. And so I think there’s a benefit to all of the opportunity we have, but there’s a trade off that we need to be able to go through and to be able to validate that the systems really behave the way we expect, and not just in, like, our test or integration environment, but in the turbulent environment of production.

Kolton Andrus: Colton is right on. Right. We’ve got a great opportunity with AI, and I think companies are really leveraging that and are doing a lot of experimentation today. The next frontier, though, is as we start seeing these things in production and how are they going to behave in production. We’re shipping code really fast. The complexity is increasing. But how can we make sure that that code is going to be reliable? How can we make sure that it’s secure? How can we govern all of that? I think it’s changing quite a bit of the landscape, and it’s also adding some questions that we didn’t have before. So from an observability standpoint, we can say things like, okay, is it running? Is it performing in the way that we expect? But with AI workloads now we have additional questions. Is it accurate? Is it working in the way that we expect it to? Is it providing answers that are accurate? And so from an observability standpoint, we need to add some of that layer to really understand how these systems are behaving in a production environment.

Swapnil Bhartiya: What kind of governance, what kind of guardrails you are also putting there from observability point of view, because we trust observatory tools that, hey, this is the metrics. I mean, data never lies, but here the risk is that, can we even trust that?

Kolton Andrus: Yeah, I think that’s one of the big differences in the approach that we take here at Dynatrace is that, of course, we leverage LLMs and we leverage genai in the platform to be able to help bring answers to the data, but we also leverage deterministic AI. And I think that’s a key component, is that you can trust deterministic answers much more and from there be able to take trusted action. And so deterministic AI is a key component of. Of observability and a foundation, I think, for, you know, leveraging then probabilistics models like LLMs and so on.

Swapnil Bhartiya: Quotren from your perspective, I mean, best practice is an overused word. But you know, what should be the right approach to kind of break that cycle of trust? And can you entrust it?

Philippe Deblois: Yeah, well, I think it’s. It’s. We share a very similar approach between Dynatrace and Gremlin here. A lot of the work that we do in the Gremlin platform is deterministic. We have a hypothesis, we’re going to go run a specific set of tests, we’re going to measure the outcome of those tests, we’re going to understand what happened, and we’re going to leverage LLMs around how we interface with the user so that they have a simplified explanation, so we can put it into plain text form. But we’re not going to leverage the LLM to process the data or to do the in depth analysis. And I think this is one of the disservices of the current trend of AI is everything has to be everything, AI has to be an LLM. But actually there’s a lot of AI that is not LLMs. And there’s this great thing called machine learning and all sorts of techniques that we’ve been honing for the last 20, 30, 40 years that can and should still be applied. And so that’s one of the approaches that we’re also taking is, hey, we’re going to look at all of this data we’ve collected across the lifetime of Gremlin. Millions of experiments, tens of thousands of different systems that we’ve analyzed and we understand the results and we’re going to distill that down into credible, actionable recommendations for customers. And then we’re going to use the LLMs for helping smooth out the interface and finding, you know, and helping to find the right APIs and find the right actions to take. But we’re going to do it in a way where we prompt and we confirm. We don’t just blindly take action and believe what’s said.

Swapnil Bhartiya: And not just LLM. It has to be the biggest LLM out there. You know, the product should be high, but that is not the case. Sometimes the tiny one can do much better job than the big ones. Now if you look at these two companies, companies, Dynatrace, Observability, Gremlin, the whole reliability, how are these solutions that you are bringing together? How should we look at it, this partnership, what impact it will have on. I’m not looking at a smaller picture of teams in general, but because AI, as I talked to Colton before, that is redefining. Every interview I am doing is like, is about AI. 99, 9.9%. Every discussion is about AI. Is it transforming the way we write code, we deploy code, we find vulnerabilities, we fix vulnerabilities. So there was a time, you know, when everybody think about software, now everything is AI. So also talk about not the smaller picture, but bigger picture when it comes to observability, reliability, AI and these two companies.

Kolton Andrus: I think, I think Colton and I are probably going to agree on this point, but I think we, you know, we’re, we’re in the business of helping our customers be proactive and Preventative. And so in other words, yes, failures will happen in production, and we have to deal with those. And from an observability standpoint, we’re watching all of that and we’re making sure that we solve problems, but we need to solve problems way earlier than that. And we need reliability testing to be able to, to help us do that. And that’s where like resilience and observability really belong together. And I think that from my perspective, the customers that we’ve helped together as a partnership have really been able to achieve some of those goals of preventing P1s from happening in the first place. It’s a lot more cost effective to be able to solve those problems early than to deal with them when they’re impacting large production environments. And if you add AI to the mix, that only compounds the problem. And so from my perspective, being proactive, starting early, doing the testing early, and using observability to augment the data that we get from Gremlin is, I think, critical for any large enterprise.

Philippe Deblois: Yeah, Philippe knows me. Yeah, 100% agree. Yeah, that’s, that’s. I’ve been out here preaching that for the last decade, trying to get folks to get in front of the problem instead of waiting for the problem to occur. I think the other thing I’d bring up is it’s an opportunity to provide more data and back to, you know, better data helps us make better decisions and helps us to build better systems and understand those systems. Well, a lot of classic observability is only what has happened, not what could happen. And Gremlin gives us the opportunity to go out and to create these scenarios and see how the system actually responds to those. Instead of looking at week over week or month over month analysis, instead of waiting for an incident to occur and deciding, hey, would that last incident be like the next incident we saw? We can go out and proactively create these experiments and these scenarios and we can understand how that impacts the system and then concretely measure that and understand what the knock on effect is. How did that impact our metrics? How did that impact our dependency? How did that impact our response time? Ultimately? How would it impact our customers? And we can leverage that to be able to go find and fix things that we might not have otherwise seen or had to wait until they had already caused customer pain to uncover and to fix.

Swapnil Bhartiya: Yeah, and if you just look at this partnership as you mentioned, what are the key benefits of this integration? What will become possible or what has become possible now, which was not possible earlier.

Kolton Andrus: Yeah. So one of the cool things that the Gremlin team did was Dynatrace is a platform and it allows customers, it allows partners to build applications in the Dynatrace platform. And so it provides all of the role access, permissions, security, the actual UI framework that’s required to look and feel exactly like the rest of the platform. And so the Gremlin team basically built an app to bring reliable reliability scores into the Dynast platform in a way that’s native to the platform, but also in a way that our existing customers already understand. And so they can be in a dashboard and they can have reliability scores from Gremlin right from within their flow and continue to use the platform in the way that they’re used to. And so it’s really about kind of bringing all of the great data that that is coming from Gremlin into the Dynamics platform in a way that users can continue to leverage in their day to day job.

Philippe Deblois: Yeah. One of my favorite sayings is if you want people to do the right thing, you need to make it easy. And so how do we make it easy to do the right thing here? Well, we know people want to do this. They’ve got many tools, they’ve got many panes of glass, they got many places they want to manage things. One of the things we’ve learned over the last five years is it’s easier to convince the engineers and the people carrying the pagers that this is good work they need to be doing so they don’t get woken up, so they don’t feel the pain. But ultimately, ultimately you need leadership buy in. Especially when you’re talking about an organizational wide problem like reliability. You can’t have one team fix reliability. It takes many hands. It takes many, it takes coordination across the entire enterprise. And so in order to do that, you need leadership involved. You need leadership to have visibility into what’s occurring. And you need to have them be able to hold people accountable and reward and incentivize the right behavior in order to accomplish that. If they already have operating procedures, they already have places where they’re looking at metrics. They built dashboards, they’re already tracking this problem. We enrich that data so that it’s easier for them to go, you know, to be able to accomplish that.

Swapnil Bhartiya: And can you talk about, you know, for those teams who are already using Dynatrace, what does day one look like after they install Gremlin app? What kind of onboarding support, kind of, you know, hand holding, the effect or that is not needed.

Kolton Andrus: The good news is that it’s quite simple. They’ve done a great job of making it super easy for, for our customers. If they’re already existing Gremlin customers, it’s as simple as going to our hub. Downloading the app and sending up the integration is very easy. From there, they’ve got, they’ve got a way to communicate with other teams in the similar language, because again, they’re using the same ui, they’re talking the same language across both platforms and can view the results together in a unified way.

Philippe Deblois: Again, we look at ways we can make it easy for people to get done what they’re trying to get done. And so we’ve done a lot of work on the Gremlin side to make it really easy to set up Dynatrace, to be able to pull in those metrics and alerts, to tie those into experiments. So similarly, when we built this app, we want to make it very seamless for people that are already customers of both to be able to quickly get valued quickly, set it up, quickly integrate, help them identify which monitors and alerts are the right ones to use, be able to pull those into their experiments so that they’ve got that safety net, they’ve got that feedback mechanism. And we’ve created ways to trigger and take action on the Gremlin items to be able to kick off the tests or scenarios from within Dynatrace. So again, just focused on making it as easy as possible for customers to get value as quickly as they can.

Swapnil Bhartiya: Is it possible for you to kind of walk us through a real world example of this integration where they actually caught a reliability risk before it turned into an outage?

Philippe Deblois: We’ve had multiple shared customers where we’ve gone through this, and of course, we just launched it live. But an example of what would happen, you know, a customer comes in, they set up Gremlin, they set up the Dynatrace integration so that they’re able to collect the data that they have, we’re able to go and identify a set of risks, or we’re able to give them the feedback about an experiment that’s run. So then you swap over. They’re in the Dynatrace app, they’re building a dashboard and they’re looking at, hey, what are the scores of my services and what are the risks that exist within my services? And so in that world, that’s where it’s very straightforward for a customer to say, hey, I came in on a Monday morning. Why do I have this new risk on my production service, what’s going on there. And then they can drill in and get details about that risk, about why it occurred, and about how to go fix it. So we’ve had quite a few examples of customers already who’ve been able to go out and either through the passive risk that we create or through the tests that are run, One of the other things that we expose in the Dynatrace app is the output of our reliability, intelligence. And so that’s when a test fails, we have a set of analysis that we run again, back to the start of the conversation, more on the machine learning side and less on the LLM side. That tells the customer, hey, here’s why it failed and here’s what to do to go fix it. So again, back to the example. On Monday morning, they come in, they saw a test failed. Why did that test fail? Well, now they can drill in and get actual feedback about why the test failed and how to go fix it. And then in the ideal scenario, they fix it, they run it again, they see that it’s passing. They’re using Dynatrace during that running experiment to understand how it’s behaving, to understand what’s going on in the system, to make sure they really groked what’s gone wrong, so they can go fix it appropriately, close that loop and end up back in a steady state where they fix that vulnerability right after it’s appeared.

Kolton Andrus: We had a large insurance company speak at our annual conference, a user conference, where we had a lot of people come to Vegas to come and listen to some of the innovations of Dynatrace. And the speaker was talking about a cloud migration that they had to do. And they wanted to be able to test for things like, for example, what would an outage do, or what would high latency bring to this particular migration. And so they were able to uncover, in partnership again with Gremlin, some risks and some P1s that they could avoid by doing this testing early. And so this migration ended up with zero outage. It was one of their big success stories. And the lady that spoke actually was promoted as a result of this. And so it’s stories like that I think that we want to continue to bring to the rest of our customer base. And let’s be proactive when you’re doing these big migrations or when you’re doing these big projects, like make sure that we do the testing early and that we catch these problems before they even make it to production.

Philippe Deblois: I love the getting promoted part of that. Story, too. That’s back to the how do we incentivize the right behavior? And it’s one of my favorite questions when I’m talking to leadership. Hey, you know, we often. We often incentivize. We often reward the folks that are great at firefighting and fixing the problems, and they’re important. But what about the team that does this work proactively and never ends up firefighting? You know, they’ve essentially prevented a whole set of failures from ever occurring. How do we recognize them? And what we need is we need proof. We need evidence. And so this gives us the opportunity to go measure those impacts, provide that evidence for folks, and allow them to go share the good news. Hey, this is the work we did. We didn’t get paged. There wasn’t an outage. Our huge migration went smoothly. We. Our app has run at five nines for the last year, and the teams can go, wait, I want to be like that team. And that’s really the behavior we want to recognize and reward.

Swapnil Bhartiya: Yeah, no, that was a great story as well. Now, if something goes wrong, do they call Dynatrace or Gremlin for this?

Kolton Andrus: Yeah, if they’re Dynatrace customers, they can definitely talk to the Dynatrace team and we’ll help them through, you know, any of the issues. Of course, you know, we work with Gremlin very closely and have partnerships at all the levels of the organization, including R and D, where we have these discussions about how we can best serve our customers and help them be successful.

Philippe Deblois: Yeah, we got the secret, you know, bad signal line, set up the shared Slack channel. So if something goes wrong, if something isn’t working right, we’re talking about it and we’re fixing it, and we’re figuring out which side or where we need to go make changes in order to accommodate things. Because, again, what we care about is making sure our customers have a smooth, great experience. And it’s nice that both teams have been dedicated there. And it’s been great working with all the Dynatrace folks over the last couple of months as we’ve been getting this ready and rolling it out anytime we have a question, anytime we’re not quite sure the right way to do something, maybe we have a little bit of feedback about ways that we could build apps better or, you know, better integrate with the platform. There’s been a great partnership. We’ve been able to smooth those out, get those things figured out and get them shipped to production.

Swapnil Bhartiya: Yeah, and that’s kind of also, I Have a lot of question and you create a perfect segue. Is that as AI adoption kind of accelerate, where do you see this partnership heading next?

Philippe Deblois: Well, I think, as we’ve discussed, there’s all sorts of opportunity to help us get in front of the wave that we’re currently riding. And as we move to different types of systems being built, we start moving from just developers augmenting the way they’re writing code to having agentic systems that are writing massive amounts of code. So we shift from CICD pipelines into agentic pipelines as we see more and more operations start to be picked up and automated. How are we going to vet that all of those are working correctly? How are we going to measure them? How are we going to provide the right safety net, the right guardrails in place? And so, to me, I think that’s the biggest potential, is we have an opportunity to have a great visibility from the gremlin perspective into our customer systems. Dynatrace is doing an excellent job measuring all the things and building their own set of great capabilities. I’m sure Filippo will talk about that. But how do we take those early signals, how do we take that analysis and really provide people that feedback loop? Because we want to be able to enable that innovation. There’s a phrase I’ve found myself saying a lot lately. We want to have our cake and eat it, too. We want to have people moving quickly and taking advantage of AI, and we want them to do that without having a lot of outages, without having a lot of defects. But that’s going to require creating these safety nets, creating these feedback loops where we’re able to do that reliability, analysis and recommendation and feed it back into those agentic systems to ensure that we’re really finding and fixing those issues and getting them, getting them fixed and getting them iterated on so that when you wake up and your code’s deployed in production, it’s reliable code, it has the right alerts, the right monitors. We know it’s working correctly. And we’ve already vetted it through a series of tests to make sure that it’s going to be able to withstand the types of issues that occur in production.

Swapnil Bhartiya: Sometimes we get carried away when it comes to AI. We do want all the agents, as much as what kind of caveats are warning you folks. Yes, it’s good. Go embrace it. But please, either have these guardrails in place, have proper culture, proper practice in place, don’t just get that everything starts looking like a nail because you have a hammer.

Kolton Andrus: Yeah, I think it kind of goes back a little bit to the earlier discussion where both Colton and I were talking about not all AI needs to be LLMs or Gen AI. I think relying on machine learning, relying on AI that’s deterministic, can be hugely helpful in making sure that we’re providing accurate answers, helping engineers as they’re trying to debug code and things like that to get the visibility they need in these much more opaque systems. Now that we see with these coding agents, agents, you know, developing on their own, with agents being autonomous and so on, having that visibility and having that explainability and auditability is really sort of the next evolution, I think, of. Of observability.

Philippe Deblois: I love that explainability that, you know, we need to hear that more. We need to talk about that more, because sometimes you, you know, step one, hey, did it do the right thing? Step two, okay, I got some hopes and dreams, but. But step three, can I prove that it did the right thing? And I think that’s really tantamount to the longevity of the systems we want to build.

Swapnil Bhartiya: Once again, thank you both, Colton, Philippe, for joining and sharing your insights, this partnership. And as you rightly mentioned, there’s so much happening. So I would love to have you folks back on the show and to talk more about, because this is a problem not going away anytime soon. We are getting more and more AI, so this field will continue to evolve. But I really appreciate your time today and look forward to chat again. Thank you.

Kolton Andrus: Thank you, thank you.

Ubiquitous AI Requires Distributed Infrastructure | Dr. Robert Blumofe, Akamai | TFiR

Previous article

Java’s Move to Monthly Security Patches: What DevOps Teams Must Prepare For | Simon Ritter, Azul | TFiR

Next article