A 90 per cent success rate sounds pretty good. If an AI agent gets nine out of every 10 tasks right, it’s easy to look at that number and see a system that’s almost ready for production. The problem starts when you ask it to do the job again. And again. And then thousands more times across the business. 

That’s becoming a more practical concern as enterprise AI agents move beyond experiments. McKinsey’s 2026 State of AI survey found that 40 per cent of respondents from organisations with more than $1 billion in annual revenue were scaling AI agents, up from 27 per cent the previous year. 

em360tech image

At that point, knowing how often individual attempts succeed isn’t enough. An agent can perform well on average and still behave inconsistently when the same work is repeated. The more work it’s allowed to perform, the more important that difference becomes. So the question for enterprise leaders is starting to change. 

It’s no longer just whether an agent can do the job. It’s whether the evidence shows that it can do that job reliably enough for the way the business actually intends to use it.

AI Agent Accuracy And Reliability Measure Different Things

There’s a fairly simple distinction hiding behind some complicated-looking benchmark terminology. AI agent accuracy tells you how often attempts succeed. AI agent reliability asks whether you can depend on that success continuing when the work is repeated. Microsoft’s ThinkingBox benchmark gives us a useful way to see the difference. 

It contains 507 workflows across business scenarios including retail, hospitality, auto insurance, banking and internal IT. Each model attempts every task 20 times, which allows the researchers to measure more than average performance. ThinkingBox uses three measures that look confusingly similar at first, particularly pass@20 and pass^20. 

They are two different metrics, and the distinction between them is important. Pass@1 measures how often individual attempts succeed. Pass@20 asks whether the agent managed to complete a task successfully at least once across 20 attempts. Pass^20 sets a much tougher standard by asking whether it succeeded every single time. 

Put simply, these measures help distinguish between an agent that can complete a task, one that succeeds often, and one that can repeat that success consistently. The difference is striking. GPT-5.4 achieved 65.36 per cent pass@1 and managed to complete 91.12 per cent of the tasks successfully at least once. 

But it completed only 25.25 per cent successfully on all 20 attempts. In other words, it demonstrated that it could solve most of the benchmark. Consistently doing so was another story. The latest ThinkingBox results make the distinction even clearer. Kimi-K3 solved 93.89 per cent of the benchmark tasks at least once, but only 13.41 per cent succeeded across all 20 attempts. 

Claude Opus 5 solved fewer tasks at least once, at 79.09 per cent, yet completed 47.53 per cent successfully every time. So which agent performed better? That depends on what you need it to do. If you’re measuring the range of tasks an agent is capable of completing, Kimi-K3 looks stronger. 

If you need the same work completed consistently, the picture changes considerably. That’s why AI agent performance and reliability need to be treated as separate signals. Capability shows that an agent can succeed. Repeatability gives us evidence about whether we can depend on it to keep succeeding.

Reliability Gets Harder To Ignore As Agent Work Scales

For a user asking an AI system a question, an inconsistent answer may be frustrating. An agent repeatedly carrying out business work creates a different problem because every new task is another point where inconsistent behaviour can become operational. This doesn’t mean a system with a 90 per cent success rate has a simple 10 per cent chance of causing a problem every time it runs. 

Real agent workloads aren’t that neat. Tasks vary, failures can be related, and an agent may take different routes to reach the same result. What average performance can’t tell us is how the system behaves under the volume and conditions the business actually expects it to handle. That gap is already showing up in production. 

IBM Research studied 20 production agent deployments and surveyed 306 practitioners across 26 domains. Reliability, which the researchers describe as consistent correct behaviour over time, emerged as the top development challenge. 

The study also found that 68 per cent of production agents execute no more than 10 steps before human intervention, while 74 per cent still depend primarily on human evaluation. That gives us a useful reality check. Organisations may be deploying agents, but many production systems are still being kept within fairly controlled boundaries. 

And the researchers found that practitioners are addressing reliability through systems-level design, rather than treating it as something a better model will automatically solve. Workflow length can add another complication. Research into long-horizon agents has found that performance can break down when tasks require longer sequences of dependent actions. 

More recent analysis of 2,518 agent trajectories found that agents often failed to recover after their first mistake and rarely identified the error themselves. The point isn’t that every enterprise agent needs to perform enormous autonomous workflows. It’s that reliability can change with the conditions under which the agent operates. 

A score collected from individual attempts can’t, on its own, tell a CIO how dependable that system will be across the workload they’re actually planning to give it.

There Is No Universal “Good” AI Agent Reliability Score

Once we accept that accuracy and reliability aren’t interchangeable, there’s a tempting next step: find the reliability percentage that counts as good enough. Unfortunately, there isn’t one. Imagine two agents with exactly the same performance. One drafts internal meeting summaries that an employee checks before sharing. 

The other changes customer account information without human approval. The first can tolerate inconsistency that would be difficult to justify in the second. The difference isn’t necessarily the model. It’s what the business has allowed the model to do. That means AI agent production readiness has to be judged against the workload:

  • How serious is a failure?
  • Will someone notice it quickly?
  • Can the action be corrected?
  • Is a person checking the result before anything happens?
  • How often will the agent perform that task?

Microsoft’s current agent evaluation guidance makes a similar distinction. It recommends aiming for an overall pass rate of 80 to 90 per cent, while saying core regression tests should approach 100 per cent consistency. It also recommends running evaluations multiple times to account for variability. 

The important part isn’t the specific percentages. They aren’t universal production thresholds. What’s useful is the acknowledgement that an acceptable average across a broad evaluation set isn’t the same thing as the consistency required from critical behaviour. NIST’s latest work on AI evaluation takes that idea further. 

Its draft TEVV-Athlon framework is designed to help organisations build assessments around their own objectives and the real-world impacts and outcomes of AI systems, including agentic systems. Rather than prescribing one universal evaluation, the framework is intended to be adapted to different applications and measurement needs. 

NIST also makes an important distinction between testing before deployment and what happens afterwards. Its March 2026 work on deployed AI monitoring notes that pre-deployment evaluations usually happen in controlled environments, while real systems encounter non-determinism, changing inputs and unexpected consequences. 

Post-deployment monitoring is therefore needed to check whether those systems continue to operate reliably in practice. So “reliable enough” isn’t really a property of an AI model on its own. It describes the relationship between the agent, the work and the conditions under which the business expects it to operate. 

That gives enterprise leaders a much more useful place to start when setting deployment criteria.

Enterprise AI Agent Acceptance Criteria Need To Reflect Real Work

The answer isn’t to throw out existing AI agent evaluation metrics and replace them with one new reliability score. Average task performance still tells us something useful. It just shouldn’t be asked to answer questions it wasn’t designed to answer. Instead, the evidence required for deployment should reflect how the agent will actually be used.

How often does it succeed?

Start with the familiar measure. How often does the agent complete the task correctly under realistic test conditions? A task success rate or pass@1 score provides a useful baseline for AI agent performance. If an agent struggles to complete the work once, there’s little value in worrying about whether it can repeat that success.

Are you enjoying the content so far?

 But passing this first test establishes capability, not production readiness on its own. The evaluation also needs to represent the work the organisation plans to give the agent. A strong average across a broad test set can still hide weaker performance on the specific tasks where failure carries greater business risk.

Does it keep succeeding?

Now repeat the work. Testing identical or equivalent tasks several times can expose variation that disappears inside an average. An agent that succeeds once and then behaves differently across later attempts presents a very different deployment proposition from one that produces the expected result consistently. 

This is where agent reliability testing adds information that ordinary accuracy metrics can’t provide. ThinkingBox demonstrates why the distinction is useful, but the principle doesn’t depend on using pass^20 specifically. Enterprises can choose the number and type of repeated tests that make sense for the work they’re evaluating. 

The aim isn’t to prove that an agent will never fail. It’s to understand how dependable its behaviour remains when success needs to be repeated.

What happens when it doesn't?

Not every failure deserves the same reliability threshold. A visible mistake that a person can correct before anything changes creates one kind of risk. A silent error that reaches a business system creates another. 

The acceptable level of inconsistency should reflect the consequence of failure, how easily the organisation can detect it and whether the outcome can be corrected. Human oversight belongs in the same calculation. If someone reviews every result before execution, the agent isn’t operating under the same conditions as one trusted to act autonomously.

How much work will it be allowed to perform?

Finally, look at the proposed deployment rather than stopping at the test environment:

  • How many tasks will the agent perform?
  • How long are its workflows?
  • How much autonomy will it have?
  • Where will people intervene?
  • Do its reliability tests resemble those conditions closely enough to tell you something useful?

Together, those questions create a more practical acceptance model:

performance + repeatability + consequence + operating conditions

None of those measures needs to replace the others. They answer different parts of the same deployment decision. The point is to stop expecting a headline accuracy figure to carry the whole thing.

Final Thoughts: Reliable Enough Depends On What You Ask The Agent To Do

Enterprise AI agents are getting more capable, but capability was never going to be the final hurdle. The more work organisations hand over to agents, the more they need to know what happens after the impressive demonstration is over and the system has to perform similar work day after day. Average performance still belongs in that decision. 

It tells leaders whether an agent can successfully perform the work often enough to deserve further consideration. What it can’t establish on its own is whether that performance will remain dependable under the volume, autonomy and consequences of a real deployment. 

That makes AI agent reliability less about chasing a perfect score and more about asking better questions of the evidence. An agent doesn’t need to be infallible to be useful. It does need to be reliable enough for the work and risk the organisation is prepared to give it. 

As agent evaluation starts to account for repeated performance and real operating conditions, enterprise acceptance criteria can do the same. And EM360Tech will continue tracking what that shift means as AI agents move from impressive capabilities to systems enterprises are expected to trust.