Don't Panic! It's Just Data 20 August 2026 29 MIN

What Is "Agent Debt" and Why Is It Breaking Your AI Code?

94% of technology leaders believe AI-generated code is higher quality when reviewed, but 78% report more incidents after deployment.

While AI-generated code, also known as vibe coding, is aiding developers in releasing software faster than ever, it seems to be simultaneously causing a greater number of production failures. 

According to research carried out by New Relic, 94% of technology leaders think that AI-generated code is of higher quality when it is reviewed, but 78% say they are experiencing more incidents after deployment. Let’s explore why this is happening.

In the recent episode of the Don't Panic! It's Just Data podcast, host Christina Stathopoulos, Founder of Dare to Data, was joined by Nic Benders, Chief Technology Strategist at New Relic to talk about "agent debt,” what happens when AI writes great code that breaks in production, and why SREs (Site Reliability Engineers) are left cleaning it up.

Benders states that the issue does not lie with AI itself but rather with a growing phenomenon that he refers to as "agent debt".

After gathering telemetry data, New Relic discovered that the core issues stem from Large Language Models (LLMs). These problems date back to the early days of ChatGPT, when the models were remarkably convincing at generating software code.

“The code looks really good, but it doesn't work in production. It creates incidents down the road,” he tells Stathopoulos. 

This is where agent debt comes in.

Also Watch: How Can AI Bridge the Gap from Observability to Understandability?

What is Agent Debt & What are its Implications?

Benders explains that Agent debt has been accumulating in the background. High-confidence, high-quality code is shipped out, but eventually software teams are bound to run into a production issue. They could also reach a point where change is implemented, but it requires ripping out huge sections of software because it has been built on an unreliable architecture.

With enterprise software teams also adopting AI coding assistants in order to speed up development, what started out as autocomplete has now developed into autonomous coding agents which can produce large parts of applications with very little human involvement.

While this has significantly improved developer productivity, it has also introduced new reliability concerns.

Unlike traditional technical debt, which is often a strategic choice to move faster, you cannot easily know the "balance" of your agent debt account.

A key issue emerging in this space is that engineers are often unable to gauge how much agent debt has built up, largely because the AI's reasoning process remains hidden from view.

As Benders noted, the balance of the agent debt account has been quietly accumulating in the background, making it increasingly difficult to trace. This isn't simply the result of poor code — LLMs generate software with a high degree of confidence, which leads developers to place greater trust in code that appears clean, readable, and complete.

It is not until after the deployment that hidden problems start to appear as incidents, leading to increased debugging complexity and necessitating costly architectural refactoring.

Also Watch: How Do You Monitor AI Agents in Production Without Breaking Incident Response?

Can Observability Aid The Agent Debt Issue?

In an era of rapidly evolving enterprise software, Benders argues that traditional development methods are no longer keeping pace.

"We need to look at observability," he explains. "It's straightforward — think of code going into production that needs to be monitored."

Yet the challenge is anything but simple. Unreliable AI architecture, flawed at its very foundation, presents a new and complex problem. When the work has been built by an AI agent, untangling those flaws becomes even more difficult — interrogating an AI agent on every decision made during the build process is, in most cases, near impossible.

“You have to approach it from this totally black box state,” Benders tells Stathopoulos. “That means observing things that we didn't think were critical before because there's something that you would just know. What services does it talk to? What's the kind of failure rate? What's the change between one version and the other?”

While these things may have once been understood intuitively, today's complex enterprise software demands a more structured approach — one where observability plays a critical role in monitoring and understanding the systems we build.

For instance, New Relic is treating AI coding agents like any other production system. They measure not just the token spend, but whether engineers are using them effectively and whether the resulting code is actually better. 

Observability has stopped being merely a tool for operations and has now become the means of validating AI-generated software over its entire lifecycle.

Rather than replacing engineers, AI can be leveraged to handle architecture reviews, production readiness checks, and security validation. This frees up experienced developers to focus on the high-stakes decisions that require human judgement, rather than getting bogged down in repetitive verification tasks.

Also Watch: Are “Vibe-Coded” Systems the Next Big Risk to Enterprise Stability?

The biggest opportunity might not be making more software, but rather, Benders says it could actually be making current software safer. The next big battle will be about data and not AI models. 

Are you enjoying the content so far?

In the future, Benders anticipates that enterprise AI will become more and more focused on the use of SRE agents which combine observability data with internal business context in order to speed up the diagnosis of incidents. The systems will not be successful simply because they are smarter; rather, they will succeed because they have access to better information.

Benders expects the major trend to shift away from the agents themselves and toward focusing on how to feed the right data into them, emphasizing that "it's more about the data than about the intelligence.”

For organisations that are rapidly taking up AI development tools, his suggestion is simply to stop judging success only on how fast AI writes code.

Rather, focus on the ways in which humans and AI collaborate throughout the entire software lifecycle, from the stages of generation and review to deployment, observability, and incident response.

In an era where AI-driven software development is accelerating at an unprecedented pace, the true competitive edge will belong not to those who generate the most code, but to those who best understand, monitor, and maintain it.

Also Watch: How Do AI and Observability Redefine Application Performance?

Takeaways

  • AI-generated code can increase production incidents despite passing code reviews.
  • 'Agent debt' is emerging as the AI-era equivalent of technical debt.
  • Observability must extend beyond applications to AI development workflows.
  • Engineering teams should measure AI value—not just token consumption.
  • AI may deliver greater value reviewing code than writing it.
  • SRE agents will compete on data quality, not model intelligence.
  • Human judgement remains essential for high-risk engineering decisions.

Chapters

  • 00:00 Introduction to the episode and guest
  • 00:30 Overview of Nick Benders and New Relic's mission
  • 01:14 Impact of AI on software development speed and review process
  • 02:55 Shift from coding to downstream operations and role changes
  • 06:31 Discrepancy between AI-rated quality and incident reports
  • 08:08 Understanding agent debt and its signs
  • 09:35 The role of observability in AI-generated code in production
  • 12:22 Measuring and understanding agent debt through observability
  • 14:10 Rethinking the AI software development lifecycle
  • 17:30 Automating code review and process checks with AI
  • 20:18 Production debugging and the role of AI in incident management
  • 21:09 Use case of SRE agents in production environments
  • 24:00 Future trends in SRE agents and data focus
  • 25:27 Key advice for using AI tools effectively in software and operations

For more information on SRE and agent debt, visit newrelic.com and em360tech.com.

The New Relic Intelligent Observability Platform helps businesses eliminate interruptions in digital experiences. New Relic is the only AI-driven platform to unify and pair telemetry data to provide clarity over your entire digital estate. We move your problem solving past proactive to predictive by processing the right data at the right time to maximize value and control costs.

Sponsored insight

Liked what Nic had to say?

Get in touch with the team at New Relic to continue the conversation.

New Relic Featured partner

Ready to put New Relic thinking to work in your stack?

Tell us about your goals. We will put you in touch with the right person on the New Relic team.

Contact New Relic