AI Infrastructure

Why AI Ships More Bugs and How to Stop Them | Kolton Andrus, Gremlin | TFiR

0

AI coding tools are shipping ten times more code and one and a half times more defects per line. Incident rates are climbing, cloud providers are going down more often, and the teams responsible for production reliability are receiving alerts only after customers are already affected. Monitoring, observability, and auto-remediation all share the same structural flaw: they act on evidence of a failure that has already occurred.

In this interview on TFiR, Kolton Andrus, CEO and Founder at Gremlin, breaks down why reactive tooling cannot keep pace with AI-accelerated development and walks through how Foresight AI uses fault injection, machine learning, and a decade of labeled experiment data to find, fix, and verify production risks before they trigger a single incident.

Guest: Kolton Andrus, CEO and Founder at Gremlin
Show: TFiR

Here is what every platform engineer and SRE needs to know.

Technical Deep Dive

Q: Is AI making distributed systems more or less reliable right now?

As of today, AI is not making systems more reliable. Teams are shipping ten times more code but generating one and a half times more defects per line, and production incident rates and cloud provider outage frequency are both increasing. AI excels at writing individual lines of code but lacks the contextual understanding needed to reason about edge cases, distributed failure modes, and the full behavior of a system in production.

“While we are in some cases shipping ten times as much code, we’re shipping one and a half times as many defects per line of code. We’re seeing more incidents, more outages.”

Kolton Andrus, CEO and Founder, Gremlin

Q: Why does AI-powered auto-remediation fall short for production systems?

Kolton Andrus, CEO and Founder at Gremlin, argues that auto-remediation is structurally reactive: it engages only after customers are already experiencing pain, and it addresses symptoms rather than root cause. Andrus notes that most Fortune 100 and Fortune 1000 customers remain reluctant to trust automated systems to make unsupervised changes to production, preferring a human review step before any fix is applied. An LLM fed raw metrics may reach an 80 percent solution, which is insufficient for systems that must be measured in nines of availability.

“We want our accuracy like we want our reliability. We want it measured in number of nines. 80% accuracy isn’t good enough.”

Kolton Andrus, CEO and Founder, Gremlin

Q: How does Foresight AI find and fix production risks before an incident occurs?

Foresight AI executes a five-step loop: it analyzes the customer system for what could go wrong, generates a targeted set of fault injection experiments, runs those experiments and evaluates outcomes, recommends concrete fixes for anything that fails, and then re-runs the original test to verify the fix resolved the root cause. This end-to-end loop closes the verification gap that most AI SRE tools leave open, because Gremlin injected the failure and therefore knows the root cause and all downstream side effects, not just the observable symptoms.

“After we fix the issue, we go back and we run the test again. We go back and we verify that the system really is in a better state. That’s the closing the loop piece that’s really missing from a lot of these automated AI systems today.”

Kolton Andrus, CEO and Founder, Gremlin

Q: How does Foresight AI surface risks that have never caused an incident?

Andrus frames the underlying compute surface as finite: regardless of what AI generates, systems still run out of CPU, memory, or disk, and dependencies still become slow, fail, or return wrong responses. By treating each system as a black box and exhausting that finite failure set, Gremlin covers the majority of issues without needing to predict application-specific bugs. The platform augments this with a library of labeled risks built from millions of past fault injection experiments and deep knowledge of misconfigurations across Kubernetes, AWS, GCP, and Azure environments.

“There’s a finite set of things that can occur, and by testing those, we’ve really covered the majority of the issues.”

Kolton Andrus, CEO and Founder, Gremlin

Q: What is the role of human oversight in Foresight AI’s automated remediation workflow?

At launch, Foresight AI produces diffs and recommended fixes for Kubernetes and AWS environments, with GitHub, BitBucket, and GitLab pull request integrations shipping shortly after. The platform also supports agentic pipelines where a customer’s own agent retrieves recommendations and applies fixes with access to source code. Most enterprise customers currently want a human to review changes before they are applied to production, but Andrus expects that posture to evolve as confidence in the system’s accuracy builds, drawing a parallel to how automated code review has gradually displaced manual review in some organizations.

“As we build trust and confidence in these systems, that’s when people will feel more confident in saying, yeah, just go ship the fix, let’s test it.”

Kolton Andrus, CEO and Founder, Gremlin

Q: How does Foresight AI function as a reliability gate inside agentic deployment pipelines?

Andrus describes Foresight AI as an adversarial agent positioned inside the deployment pipeline: after an agent has designed a feature and written the code, Foresight runs its fault injection analysis before the code ships. If the system fails the tests, that failure becomes a hard gate blocking deployment, and the findings are fed back into the pipeline so the coding agent can correct the issues. Once the code passes the full test suite, the pipeline proceeds. This approach is designed to preserve the speed benefits of AI-accelerated development without sacrificing production reliability.

“We become the adversarial agent in that situation to go ensure that the code is really production worthy, that they’ve thought about not just the happy case, but the edge cases and the failure conditions.”

Kolton Andrus, CEO and Founder, Gremlin

Q: How does Gremlin measure reliability improvements when incidents never occur?

Andrus acknowledges the attribution problem directly: if no incident happens, teams cannot easily distinguish good engineering from good luck. Gremlin addresses this through a reliability score, system baselines, and two internal personas within Foresight, a technical program manager role that tracks work in progress and prioritizes findings, and a customer success manager role that produces a record of every issue found and fixed. This output gives reliability teams a concrete artifact to present to leadership showing the value of proactive work rather than simply the absence of incidents.

“We want to go to leadership and say, look at all the things that we found and caught before they ever went out the door. This is why we’re having more reliability, less issues this year.”

Kolton Andrus, CEO and Founder, Gremlin

Q: Why is observability telemetry alone insufficient to prevent novel production failures?

Telemetry tells teams what happened; it does not tell them what could happen. Andrus points out that alert thresholds are set by looking at historical behavior, which means they are calibrated for failures already seen, not novel failure modes. At scale, incidents frequently manifest in ways that have no historical precedent, so a threshold drawn from past graphs will not catch a new failure pattern. Gremlin uses telemetry as input during experiment execution to understand system behavior under fault conditions, but treats it as a complement to proactive testing rather than a substitute for it.

“If you’re only looking at what’s happened, that isn’t going to tell you necessarily what could happen. That’s really the trap that teams fall into.”

Kolton Andrus, CEO and Founder, Gremlin

Q: What made Foresight AI technically feasible now when it was not possible two or three years ago?

Andrus traces two enabling conditions: Gremlin’s internal preparatory work over the past year to structure and label the data from millions of past fault injection experiments, and the rapid advancement of LLM capability over the last six months to a year. The LLM layer handles natural language conversation and guides analysis, but the analytical weight rests on Gremlin’s proprietary labeled dataset and agentic tooling, not on open-ended LLM inference. Gremlin validated the approach by deploying Foresight internally on its own production systems before releasing it to beta customers, several of whom told Andrus they would not go back to their previous workflow.

“Now we can take the knowledge that everyone has amassed and distill it down so that every engineer, every SRE can be like an expert.”

Kolton Andrus, CEO and Founder, Gremlin

Resources & Documentation

  • Gremlin, chaos engineering and reliability platform with Foresight AI for proactive fault injection and automated risk remediation
  • Gremlin Documentation, official docs covering fault injection experiments, Kubernetes and AWS integrations, and the Gremlin reliability score

***

👇 Click to Read Full Raw Transcript

Swapnil Bhartiya: If you look at your teams, who are of course managing the whole system, by the time they receive alert, by the time an alert goes out, we all know that the damage is already done. Monitoring, alerting, even the new AISI tools, they all react after something has happened. And by the time they speak, customers are already hurting. And it’s getting worse because now AI system ships code faster, then teams can actually react or test it. More unproven code is reaching production, which also means that things will break. And if you look at Gremlin, they have spent decades in breaking things, breaking systems on purpose, and now they are launching Foresight AI which finds, fixes and verifies risk before they become incidents. And, and once again, we have with us Colton Andrews, CEO and founder of Gremlin to talk about Foresight AI. First of all, Colton, it’s great to have you on the show again.

Kolton Andrus: Thank you. My pleasure. Always enjoy chatting with you.

Swapnil Bhartiya: It’s my pleasure. I mean, you have spent your whole career on. But you are, from the time I have known you. Before we talk about what you’re announcing today, I’m kind of curious to hear from you. From your perspective, is AI making systems more reliable or less reliable?

Kolton Andrus: Yeah, I would say not yet we’re not. It’s not making it more reliable as of today, I think the speed of execution, the, you know, quickness to deploy code, the patterns that are being matched and shipped out the door aren’t the most reliable patterns. I think AI is good at writing code. I don’t know that AI is good at building and managing distributed systems. And that’s really the name of the game. What’s in production? Is it working? Have we thought about the edge cases, not just the happy cases. And some of the metrics we’ve seen have reinforced this, that while we are in some cases shipping ten times as much code, we’re shipping one and a half times as many defects per line of code. We’re seeing more incidents, more outages, we’re seeing cloud providers and service providers go down more often or have a reduced quality of service. And so I think we’re in that period where we’re starting to feel the pain of the choices we’ve made and there’s room for improvement.

Swapnil Bhartiya: Very well, sir. Thank you. And if you look at AI powered SRE tools, if they’re actually AI powered, because these days everybody just puts the AI label in almost everything. But if we do look at AI SI tools, they’re launching everywhere, most of them are pitching auto remediation. Auto what? Do you think of that trend? Because at one point, as you’re saying, they’re not making system better. And also we all know how AI behaves, how it can hallucinate, it can do so. You have to put all the hardness and everything else. But if you just look at how the rest of the industry is approaching through auto remediation, what do you think of that? And then we’ll talk about Gremlin’s approach.

Kolton Andrus: Yeah, well, I think you nailed it on the intro is that auto remediation is after things have already broken, so you’re already feeling pain, you’re already dealing with customers that are unhappy. I think auto remediation is interesting. It’d be great if it has the degree of accuracy that we want for production systems. Most of my customers are a little gun shy about, you know, trusting those to really make production changes. They still want some review, they still want some human in the loop. And I think, yeah, my, my CTO has a comment. You know, I said, I’ve made this comment, hey, you can probably take a bunch of metrics, throw it in an LLM and get like an 8020 solution. And his quip back is, well, probably like a 2080 solution. It’s not really giving us, you know, the type of accuracy and coverage we want. And a Gremlin, what we want is we want our accuracy like we want our reliability. We want it measured in number of nines. We want 99% accuracy. 99.9% accuracy. 80% accuracy isn’t good enough.

Swapnil Bhartiya: No. So very true. So true. Thank you. Now, just look at foresight AI, which is of course built to fix issues before they impact customers. Talk a bit about the internal workings. How does that actually work?

Kolton Andrus: Yeah, so one of the things we wanted to do is just make it easy to do the right thing. As we built out Gremlin and our platform, and there’s a lot of pieces that engineers or SREs need to do in order to get the full value out of it. They need to come in, they need to analyze their system, they need to look at what could go wrong. They need to come up with a set of experiments, they need to run those experiments, they need to see how they behaved, and then they need to go find and fix the things that come out of those. And so with foresight, this is really meets the vision that we had a decade ago when we built the company. Can we just do it for you? So now we can come in and we can use both LLMs, but also machine learning and the data we’ve collected over the last 10 years, millions of fault injection experiments run, tens of thousands of distributed systems that we’ve analyzed. These are all labeled. We know how they behaved, we know what the outcome should be. We’re able to come in and we’re able to use that data to go one, look at a customer system, look at what could go wrong, to go in and suggest the set of tests that they should run to go verify those assumptions. Three, turn around and go understand what was the outcome of those experiments. Did it pass? Did it fail? If it failed, how can we fix it? And four, that’s where we can come in and we can go recommend real fixes to those issues. And since you mentioned AI sre, I think one of the things that’s fundamentally different here is SRE is really looking at symptoms of what occurred, and perhaps they’re able to go correct some of those symptoms. But are they really getting at the root cause? Because we are doing a deep analysis of the system from the beginning. We know the root cause, we injected the failure that triggered the issue, which means we understand all of the side effects. We’re able to have a much better understanding of how to go correct that issue. And then the piece that I think number five is most important is after we fix the issue, we go back and we run the test again. We go back and we verify that the system really is in a better state, that we really did fix the issue. And I think that’s the closing the loop piece that’s really missing from a lot of these automated AI systems today.

Swapnil Bhartiya: I want to look at two things as well. One is that, of course, you have explained very well, but how does foresight AI finds problem that haven’t caused an incident yet? So of course, you know, you may, of course, others, you know, they use telemetry, you did explain it, but yes, you are injecting, you know, to see what will feel. But still there, because where I’m pushing back a little bit is that in early, before AI era, it was very easy to think of some incidents. Hey, this is what may go wrong. But with AI, things can go in so many different directions. So how do you even find the cause of something that has not caused anything? Because we don’t even know what AI will screw up next.

Kolton Andrus: Well, I mean, that’s a great point because I’m sure there’s things that I can’t think of that could happen that could go off the rails. But at its root, computers are hardware. We have the same 10 things that go wrong within computing systems. I’m out of cpu, I’m out of memory, I’m out of disk. We have the network. What happens when one of my dependencies gets slow? What happens when one of my dependencies fail? What happens when we get the wrong response or something takes too long? So if we treat the systems more as a black box and we look at all the things that can happen to them, there’s a finite set of things that can occur and by testing those, we’ve really covered the majority of the issues. Now there might be some very bespoke interesting bug that lies out there and that might be where we have to go do some more in depth digging. But that’s the value of having run millions of experiments is now we’ve seen all sorts of these edge cases and conditions. We have these things called risks which are essentially linting misconfigurations. We have a really deep understanding of kubernetes. We of aws, of gcp, of Azure, of these types of environments, of the types of things people tend to get wrong and are misconfigured or cause things to behave incorrectly. And so that’s all part of the corpus we built, that’s all part of our agentic harness, the tools and the skills that go through and analyze those systems and rely upon that data we have. And we’re not just asking the LLM to guess at a bunch of things and go wrong. The LLM is great for turning what we have into plain text conversations, into being able to converse with it and being able to guide it in the direction you want. But it really relies upon the data in our platform to be able to go understand what’s most important, what’s most likely to incur. Where should I go poke on the edges to understand these failure and error conditions and go uncover and test them.

Swapnil Bhartiya: And if I’m not wrong, foresight can, you know, open a pull request with a fix, then rerun the test to prove that the fix is working. How much human is involved in that process? How much are you trusting AI there?

Kolton Andrus: Yeah, so we have it today so that it builds and creates diffs and we’re launching with Kubernetes and AWS and cloud provider integrations that do those remediations. Today we’re working on the GitHub and the BitBucket and the GitLab ones. Those will come shortly after. So that’s where we can show a pull request and go, you know, suggest a fix in the source code for it. We also view this as a place for agentic pipelines. A customer’s agent can talk to our agent to get a set of recommendations and fixes that they have access to the source code that they can go make. But, but at your heart is that question, are people really ready for automatic remediation when it comes to their production? Kubernetes clusters and I think this is an interesting topic. Some customers are going all in. Most of my customers are Fortune 100, Fortune 1000 customers and they’re definitely leaning in, but they still want a bit of that safety net and they want somebody to review the types of changes that are made or at least they want somebody to read it before they push the button that accepts it. So I think this is one that we’ll see evolve over the next year or two. It’s a bit like doing code reviews at the engineering level. I would have said five years ago there’s no way that we’re going to get rid of code reviews. Engineers are smart, they need to review and make sure the right things are, are happening. But what we’ve seen is, you know, a automated adversarial code review is often better than your average code review if an engineer isn’t really taking the time to understand what’s happening and do that real in depth analysis. And so I think we’ll get to a point where as we built trust and confidence in these systems, that’s when people will feel more confident in saying, yeah, just go ship the fix, let’s test it. And I think that’s also where the testing loop comes into play. If I have a fix, I can go push the pr, I can test it and prove that it works. Well now I’m much more likely to have that be an automated process than if I just have to take your word for it.

Swapnil Bhartiya: We are kind of moving in the direction of more and more autonomous AI agents. Can you also talk about that? While of course humans are making decisions, but are you also working on building? Of course, focus on governance, guardrails, the whole harness. So we can allow teams to let agents propose changes to production while keeping things safe.

Kolton Andrus: Yeah, so we’re a big fan of the idea of some sort of reliability guardrails. And I think that’s where this fits exactly into this process. Around these automated deploys, we think of an agentic pipeline where an agent’s gone out, designed a feature, written the code and it’s able to deploy that code out. Well, that’s a perfect opportunity to go run this type of an analysis that foresight provides, run those tests and see if it passes those tests. If it doesn’t, that should be a gate and we shouldn’t deploy that code and we should take the learnings that foresight provides and put it back into the agentic pipeline to go fix the issues. And so we become the adversarial agent in that situation to go ensure that the code is really production worthy, that they’ve thought about not just the happy case, but the edge cases and the failure conditions and we’re able to return back. Here are the ways to go correct those issues. Here are the changes you need to make in the approach or the code to go fix it. And then once we’re able to pass that gauntlet, then we feel good about pushing that code out the door. And yeah, in a fully autonomous world where that may have taken hours or days to sit in the pipeline before, maybe that all happens in minutes now and we can able to. You know, one of my favorite sayings is we want to have our cake and eat it too. How do we have the speed of moving quickly and writing ten times as much code, but not sacrifice the quality and reliability of the systems that we’re building? And I think this is one of the approaches that allows us to accomplish that. Well, that’s not a new problem for us. You know, if the tree falls in the woods, does it make any noise? You know, if an incident doesn’t occur, did we do a good job or did we get lucky? And there’s ways for us to measure that. It’s part of the reason we built a reliability score. That’s part of the reason we baseline systems. We understand the risks in them and we track the work that’s done. And so a couple of the Personas we built within foresight, we have a technical program manager and that’s part of the people part of the problem. You still have to go make sure people know what’s going on and help guide them through the work that’s being done and prioritize that. We also have what we call a customer success manager. And that’s what allows us to go measure what are the things we’ve gotten done. And maybe that’s by people and maybe that’s by agents. But let’s go see what are the issues we’ve uncovered in the system, what are the issues that we found and fixed. And let’s track those because those give us a clear indication of the success of the process, the improvements that have been made and allows us to go back and show to the business the value that has been provided through this process. We don’t want to go to leadership and say, hey, no incidents ever happened. So you know what, let’s just get rid of all our reliability teams and tools. That’s a net loss for everybody. We want to go to them and say, look at all the things that we found and caught before they ever went out the door. This is why we’re having more reliability less issues this year. This is why we’re saving money and customer pain. This is why these efforts have been very valuable in helping us to really unlock the full potential of our systems.

Swapnil Bhartiya: Can you also talk a bit about the relationship with observability and the whole observability tools? Because they do provide a lot of telemetry, a lot of data. What broke to the teams, they can learn from them. And as you also mentioned earlier, every system is a bit different though some things are common. You know, it’s all hardware at the base. But why is that data not enough to prevent the next outage? And how should observability team also look at it? And also we always talk about when something goes wrong. That’s when all the teams come together. Before that, everybody works in their own silos though. The whole idea was to break the silos, but there are silos where teams work independently. So can you talk about that aspect as well?

Kolton Andrus: Yeah, look, I think we need telemetry into our systems. We need to understand what’s happening and how they’re responding. We use those in real time to make sure things are behaving the way we expect. And we consume that data in order to have much better and more accurate recommendations and analysis of what happened. But then this is a bit the problem that we discussed at the start. If you’re only looking at what’s happened, that isn’t going to tell you necessarily what could happen. And that’s really the trap I think that teams fall into is, oh, hey, we’ve never seen this type of failure, so it’s never going to happen to us. Well, that’s often not the case. We often are bitten by novel failures, things that we haven’t seen before, things that manifest in new ways. And so we need that telemetry when it’s happening to be able to debug and diagnose it. We need that telemetry while we’re running the test to make sure we understand how the system behaves. We but it’s not going to tell us all possible things that can occur. And my favorite example is when an engineer wants to go create an alert for a graph. They look at the graph and it’s going along, and they go, great, we’ll just draw a line right above here. And if we ever trigger that line, then we know we’re in trouble. But then an incident happens and that graph goes to the moon. And all of a sudden, things look much different than they had in the past. Well, now all of a sudden, does the threshold you created good enough? Is it really protecting you? Are you cutting things off? How do you handle it when that occurs? Are you proactively able to go and protect your customers by falling back, by gracefully degrading? Failure happens. It happens quite often at scale. Not every failure needs to be a, you know, everything goes to the ground type of an issue. A lot of failures you can bulkhead and you can prevent from propagating around the system, and you could build some sort of a cache or a fallback response so that they impact customers. And that’s really what it all comes back to, is we want to protect the customer experience. We want our customers to be able to go and do whatever it is they’re trying to do with the system. And every little failure shouldn’t necessarily surface to the customer and cause them pain and frustration. We want to obscure them from that, from the customers, especially when they’re not actually stopping them from the important things they’re trying to accomplish.

Swapnil Bhartiya: This is kind of interesting question that when we look at, or when you look at foresight, AI, is it something that you always wanted to build for the early days of Gremlin and it was not possible, but now, because of AI, it is possible, or it is something that you build in response to where the market is.

Kolton Andrus: Yeah, it’s something we’ve always wanted to build. And when we started building this, we really took some stabs at this last year where we started building more reliability, intelligence, more analysis. Turns out that was great prep work because it allowed us to go start really taking the data within our system and preparing it. And that’s some of the underpinnings of what allow us to leverage the advances in LLM. But, you know, a year ago, when I went to my CTO and I said, hey, this is what we want to do, it was a little hand wavy. It was like, hey, you know, can we just go? Can we just go, you know, solve these problems for our customers? And at the time, you know, we weren’t sure how much it was really possible. And we dug in and we did our homework and we really understood what was possible. And what’s possible has changed a lot in the last six months in the last year. And we were able to really convince ourselves. We’re skeptics. You know, we sit around thinking about all day about what could go wrong and how systems fail. And so, you know, our impression of AI systems two, three years ago wasn’t that high. And we got in and we really understood what was possible, and then we used it ourselves. That’s one of the big things about how we approach Gremlin. We run Gremlin on Gremlin, we have a 5 nines available system because we are drinking our own champagne, as Jeff Bezos said back in the Amazon days. And we go through and we really test to make sure that all of our systems are reliable. It’s something that we’re all participating in. And so as we built foresight, we put it right in the hands of our engineers, right in the hands of our solution architects. And we said, go hammer on this. We want it to be accurate. We want it to be actionable. We want it to be useful. And the answer? Well, we learned some things. We found some things that weren’t quite up to par, and we went and fixed them and honed them and refined them. And some of my favorite parts have been our customers have been able to beta this for the last couple of months, and a couple of them have come back to us and said, Kolton, we wouldn’t live without this now. This changes the way we approach this problem. This is so much better than the way we were doing it before. We wouldn’t go back. And I think that gives me a lot of confidence and excitement about what is going to happen when the rest of the market is able to get their hands on this. Here you have a process chaos engineering, fault injection. It’s a little bit scary. It’s a little bit onerous. People have to go in and really do a lot of homework to understand how their systems behave. They’re not quite sure what to do. They’re not quite sure what to do with the results. They’re not quite sure how to fix the issues. Well, now we can take the knowledge that everyone has amassed and distill it down so that every engineer, every SRE can be like an expert. That’s the beauty of strong tools, is they really enable us all to do so much more and be more effective. And so now we remove a lot of that fear, that uncertainty, and that doubt, and we allow engineers to go in and be much more effective, much more quickly. Tell me what’s going on. Tell me what’s happened. Tell me what went wrong. Tell me how to fix it. Fix it for me. Great. Now I can move on. I can go back onto my day job and go build the next feature I’m building and reap the benefits without having to invest a massive amount of time or a steep learning curve.

Swapnil Bhartiya: Kolton, thank you so much for, of course, coming on the show again and talk about, of course, foresight AI, thanks for your time and as usual, I look forward to chat with you again. Thank you.

Kolton Andrus: Appreciate it as always. Swap now. Thank you.

Swapnil Bhartiya: It’s my pleasure. And for those who are watching, please, if you want to learn more about foresight AI, please go and check out Gremlin AI and we’ll see you in the next one. Thank you.

Why Multi-Agent AI Fails on Centralized Cloud | Jon Alexander, Akamai | TFiR

Previous article

Mainframe Dev Access: What Changed for Open Source | John Mertic, Linux Foundation | TFiR

Next article