Envision this: A ServiceNow ticket arrives at a big bank, and 50 customers can't log into their accounts. All the dashboards show a red indicator. About 60 to 100 people end up in a war-room meeting and remain there for about two hours. They discover the problem was a slow SQL query in a database ten hops down the stack, a system neither the login team nor their monitoring system had originally identified.
That is the situation which Anish Agarwal, CEO of Traversal, alludes to when explaining why incident response is becoming more difficult and why he believes the solution must be designed differently from the observability tools that most enterprises currently operate.
In the recent episode of the Tech Transformed podcast, Dan Twing, President and COO of Enterprise Management Associates (EMA), sat down with Agarwal, also a professor at Columbia University. They discuss how code laid out by artificial intelligence (AI) is affecting production incidents.
They also talk about what is needed to trace symptoms back to their causes within a portfolio of 5,000 applications, and how businesses decide when to switch from using AI suggestions to adopting autonomous remediation.
In terms of reliability, Agarwal says that the enterprise has to question whether the task was not just initiated but completed. However, it's also possible for a job to finish, yet the answer might not be the one you want.
How Do You Interpret Accuracy?
The Traversal CEO discerns three ways in which the development of AI has made troubleshooting more difficult. The first is that a workload can complete successfully and yet provide the wrong answer. The takeaway is that accuracy has to be assessed separately from completion. The second is stochasticity. When a system fails, it’s challenging to determine the cause since the same setup can yield different results each time it is run. Also, the number of possible combinations of models and agent-harnesses is constantly increasing. The third is comprehension. Since engineers now prompt specifications while an AI agent writes the code, the guest says, "our understanding of the code has dropped significantly".
Alluding to his own work, Agarwal talks about the enterprise control plane. In the past, execution served as a reasonable indication of a correct outcome; but with AI involved, teams now have to verify the outputs, either individually or by taking a sample, depending on the possible errors.
The Columbia University professor further explains how coding agents are trained.
“The way coding agents are optimised and constantly elevating, a lot of time is invested in compiling and running it, especially when considering the optimisation loop of the coding agents. That’s how it’s trained.”
The coding agents are finding creative ways to make the code run and execute because of the reward function. “As we think about reinforcement learning, that's the reward function: make it compile, make it run,” he adds.
The problem is that coding agents discover all sorts of ways of ignoring the real objective. The gap between the spirit and the letter of the law is growing. More thorough pre-production testing does not eliminate that gap. The CEO tells Twing that testing has probably become more comprehensive, since AI firms that provide code review now generate a greater number of tests before release. Yet a perfectly designed system can still fail when it interacts with other systems.
Often, faults occur not because your system is incorrect, but because it interacts with other systems, and failures usually happen at the points of interaction.
His analogy is of a perfect car colliding with someone at a defective traffic light. When giving the operational example, Twing referred to the situation from his work on automated teller machines. He explains that the data is passed on by an upstream system that nobody had simulated, and when the number of transactions is high, the failure spreads rapidly.
Why Move Left
When AI Runs Site Reliability
How causal AI moves SRE from dashboards to decisions, cutting MTTR and turning production telemetry into an executive resilience lever.
When Twing inquires how managing a live incident has evolved, the enterprise executive highlights a profound shift in software engineering and SRE. Driven by automated workloads and AI coding agents, the dynamics of identifying and resolving issues during an outage have changed significantly.
“LLMs are making that judgment,” Agarwal says. This is when he outlines a framework for agentic readiness levels in software reliability. He refers to an example of Traversal, stating that an enterprise customer is like autonomous vehicles: L0 represents driving manually with no assistance, whereas L5 is complete autonomy, similar to Waymo.
“You could apply a similar kind of framework for reliability, where L0 is where you're fully manually troubleshooting an incident,” he tells Twing. “L1 is where you start using runbooks. You have pre-existing runbooks that you can execute in some way, and someone can follow them. L2 is where you start using AI, using it, for example, to summarise a complex incident channel.”
According to the CEO of Traversal, customers who have been using Traversal for about six months, the coding agents are now sending requests to Traversal's MCP server to produce more resilient code before it is released into production. Some of these customers have also allowed autonomous deployment of fixes in a limited manner.
The user base has done the same. The original users were SREs and those at L1 to L3 support level. Nowadays, they are software engineers and the coding agents themselves. Twing compared the situation to the time, 15 or 20 years ago, when the industry accepted that security had to be designed in from the start.
Looking ahead, Agarwal suggests rethinking observability's underlying data primitives for agents rather than dashboards. He also recommends extending the same reasoning, "finding needles in haystacks," to security, networking and platform work.
Ultimately, enterprises need a pre-built, agent-readable model of production, a causal method for cutting through simultaneous alarms, and humans kept in the loop until trust is earned, moving stage by stage from recommendations towards autonomous remediation.
For IT leaders, the chief executive advises deciding in advance when to build in-house and when to work with a vendor, since he thinks most companies need both, and then run a pilot or proof of value that fits that strategy.
Key Takeaways
- "Did it run?" no longer proves an AI-built system worked.
- Coding agents are trained on compile-and-run, so they can sideline the real goal.
- Same configuration, three runs, three different results.
- Engineers now write specs; agents write code, so code comprehension drops.
- Testing improves, but failures still happen at the seams between systems.
- Agents need to search far more data than dashboards were built for.
- Traversal's production world model maps entities, connections and live telemetry.
- Customers are using it to validate CMDBs, which was unplanned.
- Causal analysis separates upstream and downstream anomalies from spurious ones.
- Multiple paths converging on one cause builds confidence in a root cause.
- L4 cut a war room from 60–100 people to five to 10.
- Reported MTTR fell from two or three hours to 20–30 minutes.
- Coding agents now call Traversal's MCP to write production-aware code.
- Decide up front when to build in-house versus buy.
Visit traversal.com.
Comments ( 0 )