For most of the generative AI era, productivity has been fairly easy to picture. A person has a task to complete, AI helps them do it faster, and the time saved becomes part of the business case. AI agents are starting to complicate that equation.
Instead of helping with one task at a time, agents can increasingly take on larger pieces of work and continue while the person moves on to something else. Several can also work at once. OpenAI's internal data offers an early example of how far this can go.
By June 2026, its heaviest Codex users were regularly generating more than 60 hours of agent work in a single day, spread across multiple agents running in parallel. That's an extreme example from an AI company with unusually advanced users, so it shouldn't be treated as a benchmark for ordinary enterprises.
But it exposes an important shift. A person still has the same number of hours in their working day. The amount of machine work they can set in motion no longer has the same limit. If machine working time stops mapping neatly onto human working time, measuring AI productivity mainly through individual task speed starts to tell us less about where the real gains, and constraints, sit.
AI Agents Are Breaking The Old Productivity Equation
Using AI to complete a task faster still has value. If something that once took an hour now takes 20 minutes, the saving is straightforward enough to measure. Delegation works differently. OpenAI found that by May 2026, 70.2% of sampled individual Codex users had submitted at least one task that it estimated would take a human more than an hour to complete.
More than a quarter had submitted one estimated at more than eight hours. OpenAI is careful to say these human-time estimates are model-generated and should be treated as directional rather than exact, but the broader change is clear. People are giving agents longer pieces of work rather than simply asking for individual answers.
Once several of those tasks can run at the same time, the productivity model changes again. The worker is no longer just using a tool while they work. They can delegate one task, start another, set a third running and return later to assess what each has produced. Machine execution is happening while human attention is somewhere else.
That creates a form of leverage that traditional time-saved measures don't capture particularly well. Five hours of agent work completed while someone spends one hour on something else isn't the same as making that original hour five times faster. But those five hours of machine work don't automatically become five hours of useful output either.
Someone still has to decide what work to delegate, keep track of what is happening and judge what comes back. As execution becomes easier to distribute, a larger share of productive human work starts moving upstream.
The Human Role Starts Looking More Like Management
There is something strangely familiar about this way of working. Give one agent a research task. Ask another to analyse some data. Send a third to test an approach. Check what each is doing, step in when one goes off course and decide which outputs are good enough to use. Those are closer to delegation and supervision skills than traditional tool-use skills.
That doesn't mean AI agents are employees, nor does it mean every knowledge worker is suddenly becoming a people manager. It means individual contributors can start needing capabilities that were previously associated with managing work through other people.
McKinsey has already begun describing the emergence of the "agent manager", where people increasingly coordinate AI capacity alongside human capacity. Its research raises questions about how organisational spans, roles and management models change when agents become another source of execution.
The labour market is showing a related shift. PwC's 2026 AI Jobs Barometer analysed more than one billion job adverts and found that junior roles with the highest AI exposure were seven times more likely than the least exposed junior roles to demand traditionally senior skills such as leadership and strategic thinking.
PwC also found that skills in highly AI-exposed jobs are changing more than twice as quickly. There are several forces behind those changes, so we can't attribute them all to people supervising agents. Still, they fit an emerging pattern. As machines perform more of the execution, human value can move towards judgement, direction and decision-making.
The difficult part is that those abilities don't scale simply because more AI capacity becomes available.
More Agents Don't Automatically Mean More Productivity
There is an obvious temptation here.
- If one useful agent increases a worker's capacity, surely five are better?
- If five work, why not ten?
Because execution capacity and productive capacity aren't the same thing. Every additional workstream creates something else to keep track of. A person may need to remember why the task was assigned, understand the decisions the agent made, recognise when it has gone wrong and work out whether its output still fits with everything else happening around it.
Human attention doesn't become parallel just because machine execution does. Research from Georgia Tech offers an early indication of the problem. A 2026 study involving 80 participants examined people supervising multiple AI systems in a decision-making task. When the AI systems depended on one another, participants reported greater workload and became less sensitive to some errors.
Mistakes could also become harder to recognise when one system's work affected another. This was a controlled study rather than an enterprise deployment, so it doesn't tell us how many AI agents an employee can realistically manage. In fact, we don't yet have enough evidence to set any useful universal number.
That number probably wouldn't mean much anyway. One employee may comfortably supervise several agents performing predictable, independent tasks. Another may struggle with two agents working on complex, connected decisions where a small mistake can affect everything downstream. The useful question isn't how many agents someone has running.
It's how much additional execution they can introduce before the effort of understanding and controlling it starts consuming the productivity gain. And much of that effort appears when the work comes back.
Verification Could Become The Real Productivity Bottleneck
Imagine an agent completes three hours of work in 20 minutes. On paper, the saving looks enormous. Then someone spends an hour checking its sources, correcting mistakes, rebuilding part of the analysis and making sure nothing important was missed. There is still a productivity gain. It's simply much smaller than the execution time suggests.
This creates an important distinction between machine output and useful, verified output. The more work organisations delegate, the more important that distinction becomes. An agent producing something quickly has limited value if the human responsible for using it can't establish whether it's correct without effectively doing the task again.
The verification burden also changes with the work itself. Reviewing a low-risk, familiar task may take seconds. Checking an unfamiliar technical analysis could take considerably longer. A mistake in an internal summary may be easy to correct, while an error in a financial decision, production system or regulatory process can carry much greater consequences.
So the sustainable level of AI delegation won't be identical across employees, tasks or organisations. It will depend partly on how independently an agent can work, but also on how easily its output can be checked and how much certainty the business requires before acting on it. This is where the economics of AI productivity start becoming more interesting.
As execution gets cheaper and faster, verification doesn't disappear with it. In some workflows, it could become the scarce resource. That gives enterprises a very different productivity problem to measure.
Enterprises Need To Measure Leverage, Not Agent Count
There is already strong pressure to prove that enterprise AI investment is producing returns. Deloitte's August 2026 research found that 61% of surveyed leaders expect most agents used by their organisations to become generally autonomous within four years, with humans providing oversight.
At the same time, only 5% said their business processes were highly prepared for agents, while just 15% had scaled orchestrated multi-agent adoption. Simply counting deployed agents won't close that gap. Nor will measuring machine runtime, number of tasks completed or hours supposedly saved tell leaders whether agentic AI is making people meaningfully more productive.
A better measurement lens starts with the relationship between additional useful output and the human effort required to produce it. If an employee delegates work that previously took ten hours and spends one hour directing, checking and correcting it, the leverage is easy to see. If the same employee spends eight hours supervising the process, the picture changes.
And if adding a fourth parallel workstream causes mistakes, missed context or enough additional review to wipe out the gain from the first three, more agent capacity has stopped translating into more productivity. That means AI productivity metrics increasingly need to account for what happens around execution.
- How much work is genuinely being delegated rather than shifted into another form?
- How often does the person need to intervene?
- How much time goes into reviewing and correcting outputs?
- Does increasing concurrency keep increasing useful output, or do returns begin to fall?
There probably won't be one universal formula for answering those questions. The right balance will depend on the task, the capability of the agent, the expertise of the person supervising it and the consequences of getting the answer wrong. But the underlying measurement principle is much simpler.
The aim isn't to maximise machine work. It's to increase the amount of reliable work people can produce without the cost of directing and verifying it rising just as quickly.
Final Thoughts: AI Productivity Will Depend On How Much Work Humans Can Keep Under Control
AI is beginning to loosen a constraint that has shaped knowledge work for a very long time. One person can only physically perform so many hours of work in a day. Agents can allow them to set considerably more work in motion during those same hours. That creates real productivity potential. It also moves the constraint.
When execution can scale faster than human attention, the limiting factor becomes how much of that work a person can effectively direct, understand and judge before additional machine output creates more work than value. So supervisory capacity may not become a neat new KPI sitting on an executive dashboard.
It is more useful as a way of thinking about what enterprise AI productivity actually means. The strongest agent-enabled workforce won't necessarily be the one running the most agents or generating the most machine hours. It will be the one creating the most useful, trusted output for the human effort still required.
As agentic AI moves further into everyday enterprise work, that distinction will become harder to ignore. The technology will keep expanding how much work organisations can set in motion. EM360Tech will continue examining the more difficult question that follows: whether the systems, people and measurements around it are keeping pace.
Comments ( 0 )