Buying enterprise AI is starting to create a strange evidence problem. The more capable AI agents become, the harder it is to prove what they're actually capable of doing inside a real business. An agent might need to investigate why a customer renewal is at risk, for example. 

The answer could be split across a sales opportunity in the CRM, support tickets in Zendesk, recorded calls in Gong and conversations between employees. Finding it isn't one task. It's several connected tasks, performed across systems that weren't designed to make the agent's life easy

Enterprise AI benchmarks need to test work like this if their results are going to tell buyers anything useful. But there's an obvious problem with using an actual company as the test environment. Real businesses contain customer information, employee records, commercial data and internal communications that can't simply be published alongside a benchmark. 

em360tech image

A new benchmark called Era by Eon takes an interesting approach to this problem. Instead of simplifying the test, its researchers built companies that don't exist. Each has employees, customers, sales activity, support tickets, calls, messages, documents and business applications that describe the same fictional organisation. 

Because everything is generated, the researchers also know what the correct answers should be. It creates an unusual middle ground between public benchmarks that may be too artificial and private enterprise evaluations nobody outside the organisation can verify.

Enterprise Agents Are Outgrowing Traditional Benchmarks

There's nothing particularly unusual about asking an AI model to summarise a document or answer a question from a database. Enterprise agents are being asked to do considerably more. 

McKinsey's 2026 State of AI survey found that the proportion of respondents from organisations with at least $1 billion in annual revenue scaling agents in one or more functions rose from 27 to 40 per cent in a year. Meanwhile, LangChain found that 57.3 per cent of the 1,340 professionals it surveyed had agents in production.

 As agents move into production, the work also becomes more complicated. A useful answer may depend on finding the right systems, choosing the right tools, joining information that describes the same customer in different places and working out what happened over time. This creates a problem for AI agent benchmarks built around isolated tasks. 

An agent might perform well on individual capabilities without being particularly good at combining them when a business question cuts across several systems. Enterprise teams are already feeling some of this gap. 

LangChain found that quality remained the leading barrier to production, while 89 per cent of respondents had implemented some form of agent observability. Yet only 52.4 per cent reported running offline evaluations on test sets and 37.3 per cent were using online evaluations. 

Testing the agent against realistic work would seem like the obvious answer. Unfortunately, reality comes with its own problems.

Real Company Data Creates The Opposite Problem

The easiest way to make an enterprise benchmark realistic would be to use a real enterprise. That gives the agent genuine systems, awkward data and actual business tasks to work through. It also gives the benchmark private customer information, internal files, employee data and commercially sensitive records. 

EnterpriseClawBench shows how difficult that trade-off can become. Researchers started with proprietary workplace agent sessions and turned them into 852 reproducible tasks. But because the original sessions contain internal enterprise content, the benchmark data itself can't be released. Instead, the researchers published the method used to construct and evaluate the tasks. 

Private customer evaluations have a similar limitation from a buyer's perspective. They may provide valuable evidence inside the organisation running them, but another company can't necessarily inspect the environment, repeat the test or determine whether two products were evaluated under equivalent conditions. 

So enterprise AI evaluation ends up caught between two imperfect options. Make the environment public and it becomes difficult to reproduce the complexity of a real company. Make it real and much of the evidence may need to stay private. Era tries to keep the complexity while removing the company.

A Synthetic Enterprise Can Give Evaluators Both

Era doesn't simply generate a collection of synthetic records. It generates a synthetic enterprise. The benchmark contains 83 company scenarios, covering combinations of industries and company sizes alongside 23 individually named fictional businesses. 

Its simulator fleet can represent 66 products, including Salesforce, HubSpot, Zendesk, Jira, Gong, Slack, Stripe, Amazon S3 and Google Drive. Each company uses a configured selection of them. The clever part is how those systems fit together. The Salesforce simulator doesn't invent one version of a customer while Zendesk invents another. 

Both draw from the same underlying company. A recorded Gong call can therefore involve the same contact attached to a Salesforce opportunity, while a support ticket can refer to that customer's account elsewhere. The generated companies aren't perfectly tidy either. 

They can contain duplicate customer records, inactive prospects, non-human service accounts and records left behind by employees who have already left. Activity and timestamps also follow business relationships rather than being scattered randomly through the data. 

That makes cross-system questions possible without exposing a single real customer or employee. More importantly, generating the company solves another problem that realistic data alone doesn't.

Realistic Data Is Only Useful If The Answer Is Knowable

Imagine asking an agent which salesperson has generated the most won revenue. The question sounds simple enough until the evaluator has to decide whether the answer is correct. Someone needs to calculate it. Now multiply that across different companies, systems, questions and repeated evaluations. 

Creating the answer key can become a substantial job in its own right, especially if the underlying records change. Era approaches AI evaluation ground truth, meaning the known correct answer used to grade the agent, differently. Its questions are generated from the fictional company's records and the expected answers are calculated from those same records. 

No one has to write each answer manually. The same system can also create questions that can't be answered from the information available. In those cases, recognising that the evidence is missing is the correct response. That's useful when evaluating agents because a convincing answer isn't necessarily a good one if the information required to support it doesn't exist. 

A fictional company can therefore be complicated without becoming unknowable. But knowing exactly what happened inside a fake business doesn't automatically mean that business is a good substitute for a real one.

Fake Companies Still Have To Prove They Are Realistic

Synthetic enterprises solve the privacy problem by creating a different one: benchmark realism. A fictional company can have thousands of records and still make an agent's job suspiciously easy if every customer has one perfect record, employees never leave, data is always complete and nothing ends up in the wrong place. 

Anyone who's spent much time around enterprise software may already be raising an eyebrow. Era's researchers deliberately introduced some of this untidiness. They also created a realism scorecard based on targets drawn mainly from Eon's operational data and published statistics. 

Across 23 generated companies, the mean score increased from 61.8 to 97.0 during development, while the proportion of records flagged as synthetic by its detector fell from 55.2 per cent to zero. There are still limits. 

The researchers acknowledge that generated text can't reproduce the full variety and oddities of human writing, while the realism targets themselves are only as reliable as the sources behind them. 

Era's current data plane is also read-only, so it can test agents retrieving and reasoning over information without recreating everything that happens when agents start changing enterprise systems. Its model results provide another useful warning. Nine models answered the same 33 questions three times each, with reported accuracy ranging from 42.4 to 76.8 per cent. 

Yet after statistical testing, only three of the 36 pairwise comparisons provided strong enough evidence to say one model had performed better than another. A leaderboard can look wonderfully precise. The evidence behind the ranking may be considerably less certain.

What Should Enterprises Look For In An AI Agent Benchmark?

This changes how AI procurement teams should read benchmark results. A high score isn't especially informative until you understand what the agent had to do to earn it. Before treating an enterprise agent benchmark as evidence, it's worth asking:

Are you enjoying the content so far?
  • Does the environment resemble the work the agent will actually perform? A realistic support benchmark doesn't necessarily prove an agent can handle finance or procurement workflows.
  • Do records remain consistent across systems? Cross-system evaluation becomes much less useful if each application effectively contains its own unrelated test data.
  • Is there an objectively correct answer or success condition? Someone needs a reliable way to distinguish a successful result from a plausible one.
  • Can the evaluation be repeated? Products and configurations need comparable conditions if their scores are going to be compared.
  • Does the benchmark include missing information? Enterprise agents need to recognise when the available evidence can't support an answer.
  • Are there enough tasks and repeated runs? Small score differences don't automatically prove meaningful differences in performance.
  • Does the environment reproduce real constraints? Permissions, incomplete records, messy data and restricted access can all change how an agent performs.
  • Can the agent act, or only retrieve information? A read-only benchmark tells you something useful, but it doesn't prove how safely an agent will behave when it can modify records or trigger workflows.

None of these questions makes public benchmarks useless. They make the benchmark itself part of the evidence an enterprise needs to evaluate.

Synthetic Enterprises Could Become A New Benchmark Category

Era isn't the only recent research moving towards more realistic agent environments. StartupBench, released in August 2026, builds its tasks around workflows taken from AI products with demonstrated market adoption. Even its strongest evaluated model completed only around 30 per cent of the benchmark successfully. 

Which suggests there's still a sizeable difference between performing well on conventional AI tests and reliably completing work people actually want agents to do. Researchers are experimenting with generated business environments elsewhere too. AgentMercury created 4,783 executable environments across 14 industries and 50 countries, although its purpose is different. 

Those environments are primarily designed as places where agents can learn through training rather than as public procurement benchmarks. Taken together, these approaches point towards something broader than better synthetic datasets. The environment itself is becoming part of AI evaluation. 

That could be particularly useful for enterprise agents because organisations don't operate as collections of isolated questions. They operate through people, systems, records and processes that affect one another. A benchmark that reproduces those relationships can test a different kind of capability from one that simply presents harder tasks. 

Whether synthetic enterprises become a recognised benchmark category will depend on how well researchers can prove their realism. But they could fill an important gap between simplified public tests and private evaluations that outsiders can't inspect.

Final Thoughts: Better AI Evidence May Come From Companies That Never Existed

There will probably never be a public benchmark that can tell a CIO exactly how an AI agent will behave inside their own organisation. Every enterprise has its own data, systems, permissions, processes and peculiarities. Testing against the real environment will still be necessary before an agent is trusted with consequential work

Synthetic enterprises could make the evidence available before that stage considerably more useful. They offer a way to test agents inside complicated, connected organisations without publishing anyone's actual company data, while still giving evaluators a known answer against which performance can be measured. 

The more useful procurement question may therefore be changing. Instead of asking only, "How did this agent score?", enterprise buyers can also ask, "What kind of company did it have to work inside to earn that score?" 

And as agents take on more work across enterprise systems, EM360Tech will continue examining how organisations can test their capabilities, challenge vendor claims and decide what evidence deserves their confidence.