Quantum computing has no shortage of impressive numbers. One company can announce more qubits. Another can point to higher fidelity, meaning its quantum operations are more accurate. A third can demonstrate more logical qubits, while someone else produces a record benchmark result. 

Look at any one of those achievements in isolation and the direction seems obvious: quantum computers are getting better. The problem starts when you try to work out which one is actually more capable. That question is becoming harder to avoid. McKinsey identified more than 300 organisations engaging with quantum computing worldwide in its 2026 Quantum Technology Monitor. 

Among the large global companies it analysed for spending, one-third allocated more than $10 million to quantum initiatives in 2025. As more organisations move from watching quantum development to experimenting with platforms and workloads, they need ways to judge what all those performance claims actually mean. 

em360tech image

Yet quantum computing benchmarks aren't measuring one common thing. New research from Sandia National Laboratories is trying to change that. Its Quantum Universal Operations Performance System, or QUOPS, is designed to compare computational capability across very different quantum systems. 

The fact that researchers need a new common yardstick tells us something important about the market itself. Quantum computing doesn't have a measurement shortage. It has a comparison problem.

Quantum Computers Aren't Competing On One Scoreboard

Comparing conventional computers is relatively familiar. Specifications aren't perfect, but buyers generally understand what processor speed, memory, storage and benchmark results tell them. More importantly, they're comparing systems built around broadly similar computing principles. Quantum computing is different. 

Several competing architectures are being developed at once, including superconducting circuits, trapped ions, neutral atoms, photonics and semiconductor-based approaches. Each solves the engineering problem differently, which means each can make different trade-offs between scale, accuracy, connectivity and speed. 

That makes apparently simple comparisons surprisingly difficult. A system with more qubits may have higher error rates. Another may perform operations more accurately but take longer to complete them. A third might offer better connections between its qubits, reducing the extra work needed to run a particular computation. 

These aren't small technical details around the edges of performance. They affect what the machine can actually compute. DARPA's Quantum Benchmarking Initiative illustrates the problem. Rather than assuming one architecture will win, the programme is evaluating multiple approaches to determine whether they have a credible path towards what DARPA calls utility-scale operation. 

This is where the computational value produced by a quantum computer exceeds its cost. So there isn't one scoreboard where every quantum computer moves up or down together. Before comparing scores, you need to know what each one is scoring.

The Numbers Measure Different Parts Of The Problem

None of the quantum performance metrics commonly used today is useless. Quite the opposite. They can tell researchers and potential users important things about how a system is developing. The trouble begins when one number is treated as shorthand for the capability of the whole machine.

Qubit count measures scale

Qubits are the basic units of information used by quantum computers, so counting them seems like an obvious way to compare systems. More qubits can allow a processor to tackle larger computations. But physical qubit count doesn't tell you how much reliable computation those qubits can perform

Quantum states are extremely fragile. Noise and errors can disrupt calculations, while differences in connectivity affect how easily qubits can work together. Some physical qubits may also support computation rather than being directly available to run an algorithm. A larger number can therefore describe a bigger system without proving that it's a more capable one.

Logical qubits measure something closer to usable capacity

Logical qubits move the measurement closer to reliable computation. Instead of relying on individual physical qubits, quantum error correction combines physical resources to create a more stable logical qubit. The aim is to detect and correct errors well enough for increasingly complex calculations to continue without the result becoming unusable. 

That makes logical qubits an important measure as the industry moves towards fault-tolerant quantum computing. But logical qubit count still can't tell the whole story. Two systems with the same number of logical qubits can have different error rates, operating speeds and error-correction overheads. 

Knowing how many logical qubits exist still leaves the question of what they can reliably do.

Fidelity and speed answer different questions again

Fidelity measures how accurately a quantum operation is performed. High fidelity is important because errors accumulate as computations become more complex. Speed measures another part of performance. A system needs to perform enough reliable work quickly enough for the computation to be practical. 

Even these measurements can pull in different directions. A highly accurate system isn't necessarily the fastest, while high throughput means little if errors prevent the machine from completing a useful calculation. 

The individual numbers are doing their jobs. We're creating the problem when we expect any one of them to answer a much bigger question: How powerful is this quantum computer?

One Number Hasn't Solved The Benchmark Problem

The industry has been trying to move beyond individual hardware specifications for years. Quantum Volume, introduced by IBM, combined several aspects of system performance into a single benchmark. Other approaches have followed, including Algorithmic Qubits, or #AQ, developed by IonQ, as well as measures such as CLOPS for circuit execution speed. 

These composite metrics can reveal more than a simple qubit count because they're trying to capture what happens when several parts of the system work together. But combining measurements doesn't automatically create a universal comparison. Benchmark designs make choices about what gets tested, how success is defined and which characteristics receive the most weight. 

They can also become less useful as the technology changes. IBM now describes quantum performance across scale, quality and speed, using programmable qubits, qubit operations and circuit throughput. It also acknowledges that more complex measures such as CLOPS and application-specific benchmarks can be difficult to calculate consistently from public information across different platforms. 

The problem, then, isn't that earlier quantum computing benchmarks failed. It's that the computers they're trying to describe are changing, and the questions being asked of them are changing too. That is pushing benchmarking closer to the work itself.

Quantum Benchmarking Is Moving Closer To The Workload

For an enterprise evaluating quantum computing, the most useful performance question probably isn't how impressive the hardware looks on paper. It's whether the system can produce a good enough answer to a relevant problem within a useful amount of time. We're beginning to see benchmarking move in that direction. 

IonQ published an application-focused framework in April 2026 covering 13 benchmarks across areas including optimisation, quantum chemistry, machine learning and simulation. Its main measurements are solution quality and Time-to-Solution, which measures the total time between submitting a job and receiving a result that meets a defined quality threshold. 

That distinction is important. A workload doesn't begin when a quantum processor starts running and conveniently end before anything else happens. There may be classical preprocessing, compilation, quantum execution, error handling and post-processing before the result is ready. 

Measuring the complete process gives organisations a much clearer picture of what using the system could actually involve. IonQ's framework is still vendor-developed, so it shouldn't be mistaken for a neutral industry standard. But the direction is useful. 

It shifts the question from how good are the components? towards how well does the complete system perform the task? Now another approach is trying to make that system-level capability comparable across architectures.

QUOPS Shows What A Common Yardstick Could Look Like

Sandia National Laboratories introduced QUOPS in September 2026 specifically to address limitations in existing quantum benchmarking. Rather than focusing on individual components, QUOPS tests the size of the computationally relevant quantum programmes a machine can execute successfully. 

Its QUOPS score represents the largest qualifying programme the system can reliably run, while QUOPS rate measures how quickly it can execute those units of computation. Crucially, the framework is intended to work across different hardware architectures and both physical and error-corrected logical qubits. 

Researchers have already applied it experimentally to processors from IBM, Google and Quantinuum. The results also demonstrate why better benchmarking can tell us more than another hardware record. The tested systems produced QUOPS scores ranging from 216 to 1,824. Sandia then translated the estimated requirements of selected utility-scale challenge problems into the same framework. 

Those problems would require roughly 270 million to 370 million QUOPS, leaving current systems around five orders of magnitude away in computational capability. That comparison does something a qubit count can't. 

It connects what today's machines can successfully execute with the scale of computation required for problems researchers ultimately want quantum computers to solve. QUOPS is still new. The underlying research was submitted as a preprint on 10 September 2026, so it would be premature to treat it as the benchmark the industry has been waiting for. 

Its real significance is what it's trying to make comparable: the computational reach of the whole machine.

Better Benchmarks Still Can't Tell You Whether Quantum Is Useful To You

Are you enjoying the content so far?

Even a genuinely architecture-independent benchmark wouldn't completely solve the enterprise problem. Suppose one system can reliably execute a larger quantum computation than another. That's useful information. 

It still doesn't tell a logistics company whether that system can improve its routing, or a pharmaceutical company whether it can model a molecule better than the methods already available to it. Technical capability and business usefulness aren't the same measurement. 

The relevant workload has to be considered alongside the quality of the result, how long it takes, what the computation costs and how much classical processing is still required around it. Most importantly, the quantum approach needs to be compared with the best practical alternative rather than another quantum computer alone. 

DARPA's definition of utility makes this distinction unusually clear. Its benchmark isn't simply whether a quantum system can perform an impressive calculation. The programme is investigating whether a utility-scale quantum computer can produce computational value greater than its cost. That raises the bar considerably. 

A benchmark can show that quantum hardware is progressing without proving that an enterprise has a reason to use it. The closer organisations get to real platform and workload decisions, the more important that distinction becomes.

What Enterprises Should Ask When Comparing Quantum Performance

Enterprise leaders don't need to become quantum physicists to question quantum performance claims. They do need to know what sits behind the headline number. When evaluating a benchmark or platform claim, useful questions include:

  • What is actually being measured? Is the number describing scale, accuracy, speed, computational capability or application performance?
  • Can it compare different architectures fairly? A benchmark designed around the strengths of one hardware approach may not provide a neutral comparison.
  • Does it measure a component or the complete system? Strong individual hardware metrics don't automatically translate into strong end-to-end performance.
  • What counts as success? A fast result isn't particularly useful if it doesn't meet an acceptable accuracy or solution-quality threshold.
  • Does the benchmark resemble the workload we're considering? General computational capability and performance on a relevant business problem answer different questions.
  • Can the result be reproduced independently? Transparent methodology and third-party validation make comparisons easier to trust.
  • What is the classical baseline? The relevant comparison isn't always quantum system A against quantum system B. It may be quantum against an existing classical or hybrid approach.
  • What's missing from the score? Cost, energy use, integration requirements and operational overhead can disappear from a technical performance number.

Taken together, those questions change the evaluation from “Which system has the highest score?” to something much more useful: “What does this score tell us about the decision we're trying to make?” That is a better question whether the answer comes from QUOPS, Time-to-Solution, logical qubits or whatever measurement comes next.

Final Thoughts: The Best Quantum Benchmark Depends On The Question

Quantum computing has reached an awkward stage of maturity. The machines are improving quickly enough that measuring progress is increasingly important, but differently enough that comparing that progress remains difficult. That won't be fixed by choosing one impressive number and declaring it the winner. 

The more useful shift is already beginning. Quantum benchmarking is moving beyond isolated hardware characteristics towards whole-system computational capability and, increasingly, the workloads those systems are expected to perform. For enterprise leaders, that changes how quantum performance claims should be read. 

A qubit count can tell you something. So can fidelity, logical qubits, throughput, QUOPS or Time-to-Solution. None automatically tells you whether a system is useful for the problem sitting in front of you. The best quantum computing benchmark therefore depends on the question you're asking. 

As the technology develops, EM360Tech will continue tracking how those questions change, and whether the industry's measurements are getting any closer to the answers enterprises actually need.