Artificial general intelligence has spent decades sitting somewhere between a research goal and a science-fiction idea. Now, depending on who you listen to, it could be a few years away. That sounds wonderfully precise for something the AI industry still can't agree how to define. The disagreement hasn't stopped AGI from moving into corporate strategy. 

Frontier AI labs are building towards it, investors are funding the race, governments are preparing for more capable systems, and commercial agreements are already being written around what happens if AGI is achieved. At the same time, the technology underneath those conversations is advancing quickly. 

em360tech image

Stanford's 2026 AI Index found major gains across reasoning, coding and agentic tasks, even as leading models continued to struggle with some surprisingly basic problems. For enterprises, that's where the conversation becomes more useful. The date somebody puts on AGI is ultimately a prediction. 

What AI systems become capable of doing, and whether they can do it reliably, is something organisations can actually measure.

What Is Artificial General Intelligence?

Artificial general intelligence, or AGI, generally describes an AI system capable of performing successfully across a broad range of intellectual tasks rather than being designed or trained for a limited set of them. That's the broad idea. The details become complicated remarkably quickly.

Today's artificial intelligence can already outperform humans at some individual tasks. But being exceptional at mathematics, coding or analysing documents isn't the same as having general intelligence. AGI implies an ability to apply intelligence much more broadly, including reasoning, learning, adapting and transferring knowledge between different kinds of problems.

It also doesn't automatically mean consciousness, sentience or artificial superintelligence. Those ideas are often folded into the same conversation, but they're different questions. Even the organisations developing frontier AI don't use one definition. 

OpenAI's Charter defines AGI in economic terms as highly autonomous systems that outperform humans at most economically valuable work. Google DeepMind has taken a more cognitive approach to measuring AGI progress, looking across abilities such as learning, memory, reasoning, attention and metacognition, which is essentially the ability to monitor your own thinking.

The Association for the Advancement of Artificial Intelligence (AAAI) goes back to the problem underneath both approaches: there isn't an agreed formal definition or test for AGI. So, even if a company announced tomorrow morning that it had achieved artificial general intelligence, the announcement wouldn't automatically settle the question.

Why AGI Is So Difficult to Measure

Part of the problem is that intelligence isn't one skill. A system can be brilliant at mathematical reasoning and less impressive when it needs to learn something unfamiliar. Another might remember huge amounts of information but struggle to adapt when the rules change. 

Measuring one capability tells us something useful about the system, but it doesn't necessarily tell us how general its intelligence has become. That's one reason Google DeepMind proposed a new cognitive framework for measuring AGI progress in March 2026. 

Rather than looking for one magic score, the framework identifies ten cognitive abilities that researchers can evaluate separately, including perception, learning, memory, reasoning, attention, metacognition and social cognition. DeepMind itself says there's currently a lack of empirical tools for determining how close AI systems are to AGI.

Benchmarks create another problem. They give researchers a consistent way to compare models, but successful AI benchmarks have a habit of becoming less useful as the models get better at them.

Once frontier systems can reliably solve a test, researchers need harder tests to find the limits of what those systems can do. And those limits are important. The AAAI's 2025 Presidential Panel found that researchers themselves see weaknesses in AI evaluation as a barrier to progress.

Which leaves us with an odd but important reality: AI can be improving very quickly while our ability to describe exactly how much it has improved remains imperfect.

What Current AI Progress Actually Tells Us

There isn't much doubt that frontier AI capabilities are moving quickly. The harder question is what those improvements tell us about progress towards AGI. Right now, the answer depends heavily on what you measure.

Benchmark Scores Are Rising Faster Than Generality

Stanford's 2026 AI Index describes what researchers call the "jagged frontier" of AI particularly well. Google DeepMind's Gemini Deep Think reached gold-medal performance at the International Mathematical Olympiad. Yet the strongest model Stanford evaluated for reading analogue clocks managed just 50.1 per cent accuracy. 

AI agents improved from around 12 per cent to 66 per cent success on OSWorld, a benchmark based on real computer tasks, but still failed roughly one in three attempts. ARC Prize offers an even clearer example of why benchmark performance and generality shouldn't be confused.

Claude Opus 5 scored 97.5 per cent on ARC-AGI-1 and 90.4 per cent on ARC-AGI-2 at maximum reasoning effort in July 2026. On the newer ARC-AGI-3 evaluation, its high-reasoning variant scored 30.2 per cent. ARC-AGI-3 deliberately changes the challenge. 

Instead of giving AI another collection of familiar-looking problems to solve, it places systems into unfamiliar interactive environments where they have to work out the rules, learn from what happens and adapt their behaviour.

The lower score doesn't erase the progress demonstrated by the earlier results. It tells us something different about where the limits currently sit.

AI Is Becoming More Autonomous Too

Benchmarks aren't the only place those limits are moving.

METR tracks what it calls the task-completion time horizon of frontier AI agents. Put simply, it looks at how difficult a task an AI can complete reliably, using the time a human expert would need to complete the same task as a measure of difficulty. Its current evaluation uses more than 100 software tasks.

It's an important distinction because a longer task horizon doesn't mean an AI can simply be left alone to work reliably for that many hours. What it does give researchers is a way to track whether agents are becoming capable of handling increasingly complicated work. The UK's AI Security Institute is seeing similar progress

Its frontier evaluations found that the best models now complete apprentice-level cyber tasks around half the time, compared with less than nine per cent in late 2023. In 2025, its researchers tested the first model able to complete any expert-level cyber tasks. So autonomy is increasing. Reasoning is improving. Models are handling harder tasks and working across more domains.

But those capabilities aren't advancing at exactly the same rate, which makes squeezing all of that progress into one AGI countdown rather difficult.

Why AGI Timelines Still Tell Us Less Than They Seem

Ask when AGI will arrive and you can get very different answers from people with very good reasons to understand the technology. At the World Economic Forum in January 2026, Google DeepMind CEO Demis Hassabis estimated that AGI remained five to ten years away. He said the path was becoming clearer, but important scientific ingredients were still missing. 

Leaders at Anthropic and OpenAI have discussed much shorter timelines, including the possibility of highly capable systems around 2026 or 2027. Anthropic, for example, told the US government in 2025 that it expected what it calls "powerful AI" could emerge as soon as late 2026 or 2027. 

Its definition includes systems capable of Nobel Prize-level intellectual work across multiple disciplines and independently reasoning through complex tasks over hours, days or weeks. Researchers aren't convinced that getting there is simply a matter of making today's systems bigger. 

In the AAAI Presidential Panel's survey of 475 members of the research community, 76 per cent said scaling current AI approaches to produce AGI was unlikely or very unlikely to succeed. Those positions don't necessarily contradict each other. 

People are working with different definitions of AGI, different assumptions about how AI will develop and different ideas about which technical breakthroughs are still needed. An AGI timeline can tell you what somebody expects to happen. It can't tell an enterprise what its technology will be capable of when it does.

AGI Could Matter Before Anyone Agrees It Has Arrived

There's already a rather unusual example of what happens when a fuzzy technical concept meets a very concrete business agreement. Microsoft and OpenAI's revised partnership includes specific provisions for AGI. If OpenAI declares that it has achieved AGI, an independent expert panel must verify the declaration. 

The agreement also sets out what happens to Microsoft's intellectual property rights around and after that point. Think about what that means for a moment. The AI research community doesn't have one universally accepted test for artificial general intelligence, but a commercial agreement worth billions of dollars already needs a process for deciding whether it has happened.

Are you enjoying the content so far?

Enterprises don't need contracts quite that dramatic for the same underlying issue to reach them. If AI systems become capable of completing much broader work with less supervision, that can change automation strategies. If they can transfer knowledge between unfamiliar problems, assumptions about which work requires specialised systems may change. 

Improvements in reliable autonomy could alter workforce planning, procurement and the amount of responsibility organisations are comfortable handing to AI. None of those decisions needs to wait for everyone to agree that we've reached AGI. The more useful question becomes: What can this system now do reliably that it couldn't do before?

What Enterprise Leaders Should Watch Instead of an AGI Date

For enterprise leaders, tracking AGI progress doesn't need to mean deciding whose prediction is right. There are four more practical capability signals worth watching together.

  1. Breadth. Look at whether a system performs strongly across genuinely different types of intellectual work. Being exceptional across several variations of the same task is useful, but it isn't the same as demonstrating broad intelligence.
  2. Adaptation. Watch what happens when the system encounters something unfamiliar. Can it work out new rules, learn from experience and transfer existing knowledge into a problem it wasn't specifically prepared to solve? This is one reason newer evaluations such as ARC-AGI-3 are useful.
  3. Reliable autonomy. Capability becomes more significant when an AI can maintain performance across a longer chain of work without needing someone to repeatedly correct, redirect or rescue it. The word "reliable" is doing a lot of work here. Completing a complex task occasionally and completing it consistently create very different enterprise use cases.
  4. Real-world reliability. Benchmarks are controlled by design. Businesses aren't. Enterprise environments contain incomplete information, changing requirements, legacy systems, unusual exceptions and people who don't always behave the way a test expects them to. A capability becomes much more useful when it survives those conditions.

None of these signals gives us an AGI test. Nor should an enterprise invent one simply because the wider industry hasn't agreed on one.

What they provide is a better way to interpret the next major model release. Instead of asking whether a benchmark score proves we're closer to AGI, leaders can ask which underlying capability improved, how broadly the improvement applies and whether there's evidence that it holds outside the test.

Final Thoughts: AGI Will Matter Before the Label Does

Artificial general intelligence is often discussed as though we'll eventually cross a clean technical line. On one side we'll have AI. On the other, AGI. The evidence we're seeing today suggests the journey may be considerably messier. Reasoning can improve faster than adaptation. 

Models can become more autonomous while remaining unreliable in unfamiliar situations. A system can reach extraordinary performance on one benchmark and struggle when the test changes. All of those capabilities can continue advancing while researchers disagree about whether the combination deserves to be called AGI.

Enterprises don't need to resolve that disagreement. They do need to recognise when changes in AI capability are significant enough to change their own assumptions about automation, technology investment and the work these systems can reliably perform. So when the next AGI prediction arrives, the useful question isn't whether the date sounds convincing. 

It's what has actually changed. And EM360Tech will continue following the research, frontier models and enterprise consequences behind those changes, because the most important shift may not arrive with an AGI announcement at all. We may eventually look back and realise that AI crossed the thresholds organisations cared about one capability at a time.