Cloud Native

Born Observable: Fix AI Code Before It Breaks Prod | Greg Leffler, Splunk | TFiR

0

AI coding assistants ship functions, services, and entire applications with no instrumentation attached. When those apps hit production and break, teams have no traces, no spans, and no business-logic context to hand a remediation agent. Bolting on observability after the first outage is not a strategy; it is a recovery tax paid in downtime and engineering hours.

In this interview on TFiR, Greg Leffler, Director of Developer Evangelism for Observability at Splunk, walks through how the born observable approach, OpenTelemetry-native tooling, intelligent data tiering, and the Cisco and Splunk unified data fabric give platform teams the signal they need before production breaks, not after.

Guest: Greg Leffler, Director of Developer Evangelism at Splunk
Show: TFiR

Here is what every platform engineer and SRE needs to know.

Technical Deep Dive

Q: How are developers reacting to AI-generated code, and what new risks does it introduce?

Greg Leffler, Director of Developer Evangelism for Observability at Splunk, says developers are broadly split between excitement at the speed AI unlocks and concern about what the shift costs them professionally. The deeper technical risk is that AI accelerates code volume without adding instrumentation, so teams lose visibility into what is actually running. Leffler notes that writing the code was never the hard part: knowing what to write, architecting for scale, and ensuring it is observable in production are where the real challenges begin.

“The code was really never the hard part. It was knowing what to write and then making sure that would work in the actual business you needed it to work in.”

Greg Leffler, Director of Developer Evangelism, Splunk

Q: What does “born observable” mean and how does Splunk Observability Studio implement it?

Born observable means instrumentation is added during development, inside the IDE, not retrofitted after a production incident. Splunk’s Observability Studio integrates directly into the developer’s editor, instruments code as it is written, and can audit existing codebases to identify what is and is not observable. Critically, it can instrument around business logic, recognizing patterns such as the start of a checkout transaction and adding the appropriate spans automatically.

“It can instrument around business logic. It can say this looks like the start of a checkout transaction and instrument that for you.”

Greg Leffler, Director of Developer Evangelism, Splunk

Q: Is Splunk Observability Studio open source, and which IDEs does it support?

Observability Studio is open source. It ships as a VSIX extension, which covers most major IDEs, and Splunk is actively encouraging the community to extend it to additional platforms and other observability vendors. Leffler emphasizes that the instrumentation it produces is OpenTelemetry-native, not proprietary to Splunk, so teams avoid vendor lock-in from day one.

“It’s open source and we’re hoping that people will start to expand it to other IDEs and that it will integrate with other observability vendors.”

Greg Leffler, Director of Developer Evangelism, Splunk

Q: How does OpenTelemetry help control telemetry volume and cloud observability costs?

OpenTelemetry provides built-in mechanisms to reduce data volume: tiering (hot data to your observability platform, cold data to cheap object storage like S3 Glacier with federated search on demand), sampling, and span dropping for low-value flows. Leffler also points to Splunk’s Ingest Processor and Edge Processor, which filter and transform data at the collection layer before it reaches the online store. The key principle is owning your own pipeline so that routing decisions can change without re-instrumenting applications.

“You should own your pipeline. Where you send it and how you configure that, because it’s your data at the end of the day.”

Greg Leffler, Director of Developer Evangelism, Splunk

Q: What specific cost-reduction strategies work for AI-generated telemetry volumes?

Leffler recommends deciding before deployment what data the application should emit and at what log level, then actively dropping spans for flows that do not justify the cost of online storage. He frames this as a craft decision: elegant telemetry that is cost-effective and targeted is itself a software engineering skill. He also flags that Splunk is introducing new pricing for observability-oriented logging use cases to give teams more flexibility.

“Some of the data you’re just not going to want to or be able to look at, given how much it costs to process. I do think it’s the most realistic answer.”

Greg Leffler, Director of Developer Evangelism, Splunk

Q: Why does full-stack observability break down during real incidents when app and network data are siloed?

Leffler draws on years as an SRE to describe the pattern: the first instinct in any incident is to rule out the network, which historically meant waiting on a separate team to check separately. That delay costs time, strains relationships between teams, and blocks remediation. Siloed telemetry means app teams cannot rule out network causes themselves, and network teams cannot see application context, so both sides are guessing.

“Integrating those two things has been the brass ring for my entire career. People have said we want to be able to see everything in one place.”

Greg Leffler, Director of Developer Evangelism, Splunk

Q: How does Cisco Cloud Control close the gap between application and network observability?

Cisco Cloud Control aggregates data from compute, network infrastructure (Meraki, Catalyst), wireless, end-user telemetry (ThousandEyes), collaboration (WebEx), and application observability (Splunk) into a single interface. Leffler acknowledges this unified view is not fully realized today but describes it as the active direction, with meaningful portions already functional. The goal is for a remediation agent to be able to traverse from application signal to infrastructure root cause without a human handoff between teams.

“Cisco is probably one of the only places that actually will have all of that data available and will make it usable in one place.”

Greg Leffler, Director of Developer Evangelism, Splunk

Q: What is the first practical step for a team still bolting on observability as an afterthought?

Leffler is direct: adopt OpenTelemetry first. Teams still emitting Jaeger or Zipkin traces need to migrate before anything else, because the entire ecosystem of tooling, including automatic injection and AI-assisted instrumentation, requires an OTEL foundation. Once on OTEL, the OpenTelemetry Injector can automatically instrument services on restart with no code changes required.

“If you’re not using Otel, you can’t benefit from that. Adopting OpenTelemetry as the first step means everything else you do beyond that becomes a lot easier.”

Greg Leffler, Director of Developer Evangelism, Splunk

Q: After adopting OpenTelemetry, how should teams prioritize what to instrument first?

Leffler recommends mapping the business-critical path: the flows that directly generate revenue or deliver the core customer experience. Teams should instrument the payment processing service before the Christmas-logo-swap service. Starting at the highest-impact point builds the organizational case for expanding instrumentation, justifies headcount and AI tooling spend, and produces the data that management can actually see the value of.

“How do you instrument the world? You start with the thing that has the most business impact, which may not be the easiest one or the most fun one, but that’s where you’re going to start to get the value.”

Greg Leffler, Director of Developer Evangelism, Splunk

Resources & Documentation

***

👇 Click to Read Full Raw Transcript

Swapnil Bhartiya: Hi, this is Swapnil Bhartiya and we are here at Splunk and today we have with us Greg Leffler, Director of Developer Evangelism for Observability at Splunk. Greg, first of all, great to have you on the show here.

Greg Leffler: Thank you. I’m really glad to be here. I’m looking forward to it.

Swapnil Bhartiya: How do you, as well as developers, they perceive AI and the whole software development?

Greg Leffler: I think the people are mostly of two minds that I talk to, and one is, hey, this is great. We can develop so much faster, we can get things done quicker, we can get stuff that even I’m not super familiar with. I can still develop it, I can still push it out, I can still make it work. And it is really, it’s one of the best times, I think, in history to be a software developer is because you can explore and you can iterate so fast and you don’t need to spend a lot of time reading through hundreds of pages of docs or trying to read FAQs or getting into a chat channel with people. You can just ask your agent to do it and build the code. So it’s exciting from that perspective. But I think people are perceiving sort of a sacrifice to the craft of software engineering and the effort that you spend and the polish and the individual personalized touches. It’s something that people take pride in this job, and I think that’s something that we need to figure out how to reconcile with AI. How do you make sure that, as you were saying in the beginning, I wrote this app and I made sure that it works and it was tested and it was secure and it did the things we wanted and is testing now. The thing that is the craft is it documentation, is it marketing? I think what it means to be a good software engineer is definitely going to change, has changed, and I don’t think it’s going to go back. So I think people are concerned about that. And then I think people are also concerned about is AI going to take my job? And it’s like, am I still going to have a job next week? And what is my job going to look like? Writing prompts is a very different job than writing code and being able to know how to write a good prompt and how to optimize that and how to deal with all of the restrictions around spending and around security and around where you can deploy it. And then dealing with the telemetry and the observability aspect. The job is becoming a lot harder. I think you would say it’s easy now because AI writes the code. The code was really never the hard part. And maybe people are going to come at me for that. But I think the writing the code really wasn’t the challenging part. It was knowing what to write and then making sure that that would work in the actual business you needed it to work in. If you’re writing something for five users, that’s very different than writing something for 5 million. And so figuring out how to scale and architect that is going to be where the challenges start to come.

Swapnil Bhartiya: I think now the problem is also new, different kind of problem is that because the amount of code generated. Now the second phase is that application, something goes wrong. Now, since you did not bake telemetry from the beginning, now you’re bolting it on, monitoring, bolting it on after something happened. By that time, it’s too late. So if you look at Splunk or its approach of born observable, what does it mean and how does it also, once again, takes developers back to what they enjoy doing.

Greg Leffler: Part of the hassle of instrumentation is that it’s not fun, it’s not interesting, there’s not a cool instrumentation. It’s like you stay up all night figuring it out. It’s a slog. And that’s what AI is good for. AI is good for dealing with slogs. And Splunk came up with a tool called Observability Studio that integrates in your IDE as you build your applications, it instruments them as you go, and it can even deal with existing code bases. And there’s skills that we provide that you can just say, like, audit this application. Tell me what is observable, tell me what isn’t. And then again, since it’s just an AI, you can then say, fix it. And we have a tool that can go through and add the instrumentation for you. And it can even do what previously was purely manual. It can instrument around business logic. So it can say like this looks like the start of a checkout transaction. And it can instrument that for you and give you that information. So I think it’s really going to help people develop apps the right way. I guess it’s an opinionated statement, but in a way that you can be proud to operate. And when something goes wrong, because it’s not if it’s when something goes wrong, you can know that you’ll be able to fix it. More importantly, if you have all the data available already, you can have your agents fix it and you don’t have to worry about, oh, now it’s broken and I don’t have any data and I don’t know how to instrument it and I don’t know where I need to instrument. You can just tell your agent, hey, go look at my data and figure out what’s wrong. So I think it’s a neat approach to try to have computers do what computers are best at, which in this case is tedious, repetitive work and like going through and manually instrumenting every function. That’s exactly what we should be using for. So I’m pretty stoked about Observability Studio. I actually think it’s a pretty neat thing that we’ve developed and we’re making it. It’s open source and so we’re hoping that people will start to expand it to other IDEs. We covered most of the main ones. There’s a VSIX, so pretty much everything. But we’re hopeful that people will use it for other platforms and that will integrate with other observability vendors. Because one of the things you can do is run a local OpenTelemetry collector on your machine so that even if you don’t have an observability cloud service, you can still instrument your apps, you can still see that the data would go so that when you do have that service, you can use it, or you can sign up for a free Splunk cloud and you can use that. There are options. But I think the cool part is we want everybody to instrument their applications from the beginning and to do it with OpenTelemetry. It’s not like we’re locking you into a Splunk thing, right? It’s OpenTelemetry. It’s the industry standard. And so I think it’s pretty cool. I mean, I can geek out about OTEL for a long time, but I think it’s actually pretty nice because you do get to see as it’s instrumenting your application like it tracks all your business logic, it tells you what’s going on and so you can see where the problems could be.

Swapnil Bhartiya: And since you brought in open source and of course you mentioned OpenTelemetry, let’s talk about telemetry also, because when it comes to telemetry volumes and cloud cost, they’re kind of exploding. And talk a bit about how is Splunk kind of using open standards, open technology like OpenTelemetry to help teams maintain the whole high fidelity debugging without breaking the bank, getting to the whole sinkhole of let’s not even talk about tokenomics.

Greg Leffler: Thank goodness. Yeah. Splunk, we’ve been a contributor to OpenTelemetry for a long time and we’re one of the leading contributors. It’s in our DNA at this point. And OTEL has a bunch of features built in to help reduce data volume and reduce costs. And there’s a couple of strategies. Probably the lowest hanging fruit. One is tiering, where you just say some of this data we need access to right away. It’s our online performance data for what’s going on in the application right now. And so that stuff you want to send to your observability platform, you want to pay the high cost for that because you need to access it right away. But a ton of data that your apps generate realistically is probably never going to get read, right? Like there’s audit logs, there’s security stuff, there’s debug logging if you have that turned on. And so you can send that somewhere cheaper and you could send it to an S3 glacier type of bucket, and you could use federated search and Splunk to only pull that in when you need to. And that’s one way to save money. Another approach is to use tools like every vendor has one, but like Splunk has ingest processor and Edge processor, which run where you send your data and can then say like, we’re going to sample, which you can always do. You can also just drop some data points if they’re from an application that isn’t super relevant, or you can drop most of your successes. I mean, there’s a lot of strategies of just you don’t actually need all of the data that gets generated. What you need is the right data and you need to make sure that you can go back and get anything at any point in time, right? So like, if you start to experience some weird bug, it’s really important that you have the capability to say, send everything into my online platform. So you want to have a pull cord of something’s going on. And I want to make sure that everything gets sent online. So that’s probably one of the keys is OTEL lets you decide how you build your data pipeline. And you can send the right thing to the right place at the right time. And then once you’ve built that pipeline, you can move that around to any of your other applications. If you switch vendors, you can keep the same pipeline, you can keep the same processing. And so part of it is your vendor, like Splunk should have some tools for you that help with that. But ultimately you should own your pipeline. And where you send it and how you configure that because it’s your data at the end of the day and you know better than anybody what of it is worthwhile and what’s not at any given point in time. And then I’ll also say that, like, the AI telemetry volume problem is not going to get any better. And so, like, really being thoughtful before you deploy the application of like, what data is this emitting? What level are we logging at? Like, can we tone that down? Like, do we have traces for every single flow in this app? Do we need to actually emit all of those? Again, you want them instrumented. But it makes a lot more sense to me to say we’re going to drop some of these spans, we’re going to drop some of this data because it’s just too expensive to send it to an online store. Right. And so it’s not a super satisfying answer to say just don’t look at some of the data. But I do think it’s the most realistic answer is, like, some of the data you’re just not going to want to or be able to look at, given how much it costs to process. Splunk, we’ve also introduced, we’re introducing new pricing for logging for observability use cases. Obviously, we want people to be able to analyze stuff and to look at things in the same interface. And so we’ve done some work on that. But it’s always going to cost money, right? Like, it’s always going to be something we have to make money. There’s always going to be a cost. So figuring out what you actually need to get value from is a decision you have to make. Like, it’s a taste element of, like, do we need this? Do we not? And it could be another aspect of the craft of software engineering is emitting elegant data and having elegant telemetry and having it be reasonable and cost effective for what you’re doing.

Swapnil Bhartiya: And you know, for a long time, the full stack. Observability has been an industry buzzword for years. But what happens is during a real outage, when app teams and network engineers, they are trying to figure out they hit a wall where everything, the telemetry is siloed, fragmented. Talk a bit about how when you look at Cisco and Splunk, combining application context with network intelligence, closing that gap.

Greg Leffler: Yeah, I mean, I was an SRE for a really long time and I can’t tell you, probably hundreds of incidents where I wouldn’t even start looking until I could make sure it wasn’t the network. I would immediately blame the network. It’s something wrong with the network. It’s DNS. It’s something in the network is the problem. The way that I had to do that was to call while I was on IRC and I had to chat somebody on the network team and say, could you look at these things in this network? And then I had to wait for them to get around and look at it. Most of the time it wasn’t the network’s problem. The networks are pretty reliable, knock on wood for that one. But generally speaking, the network wasn’t the problem. Most of the time is DNS the problem. But now if I want to do that troubleshooting, I don’t want to wait to engage the network team. I also don’t want to waste the network team’s time. I don’t want to call them every time something goes wrong. I want to be able to look at that myself. Ideally, I want the system to tell me, hey, we saw this outage happen. We kicked off a remediation agent, it ran through your infrastructure and found out that, yes, this switch failed, or it’s a bug in your application code, or a backhoe ran over a fiber cable somewhere. Integrating those two things has been, I mean, it’s been like the brass ring for my entire career. People have said we want to be able to see everything in one place. For a while, this was network and application data, and then it was application and security data and than it was infrastructure data. And I do think Cisco’s one of the only people that can credibly say you can do all of those things. With Cisco Cloud Control, you can look at everything from your compute to the network it runs on, to your end users, to the wireless infrastructure in the offices, to the applications you’re running. With Splunk, it is all in one place. Now, are we there today where you can actually troubleshoot? No. But it is the vision. It is coming. There are big parts of it that you can do right now. And I do think Cisco is probably one of the only places that actually will have all of that data available and will make it usable in one place. We’re really ahead of the curve on our adoption with this and we keep showing the same. There’s a lot of stuff that we can do in Cisco Cloud Control that actually will make it so that we don’t have to blame other teams. And it helps with people becoming better and more adept at troubleshooting and at solving problems. If all you’ve ever troubleshot is your application’s code, that’s all you can troubleshoot and you can’t grow your skills. And we’re going to have a real problem with not having people that know what they’re doing in troubleshooting if we never give them exposure to it. I think having all of the data in Cisco data fabric, being able to look at it in Cisco Cloud Control and then having the data from Meraki and Catalyst and WebEx and Splunk and ThousandEyes and all of those things all in one place means you can confidently say we know this is where the problem is faster. And that’s always been the dream. You know, you’re an SRE, you want to get back to sleep, right? It’s two in the morning, something broke, like you want to fix it and you want to go to bed. So having all that in one place and actually being able to use it, I think is something that is coming.

Swapnil Bhartiya: No, so very true. And also Cisco in a unique position because the network or hardware stack that no other company would love to be in there. So you do have that advantage as well. Now let’s look at a team that is still doing observability the old way, where they’re bolting things on as an afterthought. What is the first practical step you would suggest them to take as they start a shift towards born observable mindset?

Greg Leffler: Well, I think it’s the same advice I would give to most people, which is like, you’ve got to adopt OpenTelemetry as the first step that you take. Right. So if you’re still emitting like Jaeger data or Zipkin traces or whatever, you need to migrate to OpenTelemetry. And taking that first step means that everything else you do beyond that, it becomes a lot easier. Right. We’ve got a lot of tooling in OTEL to help reduce the burden of instrumentation. So even if you’re not using AI, we have the OpenTelemetry Injector. So you can install a component on the same machine and then restart your service and it automatically gets instrumented. That’s something that we provide with OpenTelemetry. If you’re not using Otel, you can’t benefit from that. I could say the classic answer of do an audit of all of your functions and all of your applications and figure out what talks to what and where it has the most customer impact. But realistically, it’s like, adopt OpenTelemetry is the first step. The second step is probably consider business context and figure out what parts of your application are on the critical path for your company making money or your customers having a good experience or whatever your North Star is. A lot of people don’t think about apps in that way, but businesses run applications to make money. The utility providers run applications to keep their customers happy, to make money. But your application, there’s a reason it exists and you need to understand what that is as an observability person and make sure that you’re measuring stuff and that you start instrumenting what is on that business critical path. Right. And so you don’t need to instrument the application that changes your logo on Christmas to have a Santa hat. Like nobody cares about that. What you need to instrument first is the application that takes credit card numbers and charges customers, that needs to be instrumented first. So it’s two parts, I guess. One is like, if you’re not using OTEL, start using OTEL. The second is, how do you instrument the world? You do it a drop at a time. But it’s like, how do you instrument the world? You start with the thing that has the most business impact, which may not be the easiest one or the most fun one, but that’s where you’re going to start to get the value. And that’s where you get your management to let you start instrumenting more. That’s where you can say, we need another person to help us with this. We need more AI credits to help us do this. You need to show value and so that’s where you start to show value.

Swapnil Bhartiya: Greg, thank you so much. It was a great conversation where you do touch upon some of the pain points that teams feel, but also the solution technologies, Cisco and Splunk are building to once again let teams, and there are like so many teams, be able to do what they want to do without worrying about all those complexities. Once again, thank you for your time and I would love to have you back on the show.

Greg Leffler: Awesome. Great. I’d love to be back. Thank you very much.

Swapnil Bhartiya: Thank you.

AI Inference Costs: Growth vs. Margin Trade-off | Ari Weil, Akamai | TFiR

Previous article