Key takeaways:
- Three papers demonstrate quantum advantage with built-in validation, enabling trustworthy quantum computations beyond exact classical verification.
- IBM and UChicago used doped Clifford sampling and spacetime codes to certify classically hard quantum computations.
- Qedma, RIKEN, and BlueQubit observed quantum phenomena beyond leading classical simulations using validated error-mitigation techniques.
- Algorithmiq demonstrated how trusted quantum results can emerge by validating the computation process rather than a classical answer.
- The Quantum Advantage Tracker continues to pit quantum advantage claims against the best available classical methods.
- Quantum computing is entering an era where beyond-classical results can be produced with rigorous evidence of reliability.
A quantum advantage occurs when a quantum computer performs a computation beyond what classical computing can achieve alone—and when the result can be rigorously validated. But this raises a fundamental question: how can we trust the output of a quantum computer when classical verification is no longer available?
Today, a trio of papers from researchers at UChicago, Qedma, and Algorithmiq in collaboration with IBM are reporting demonstrations of quantum advantage, based on frameworks designed to build trust in the quantum computation.
Since the emergence of quantum computing, we’ve always relied on a classical computer to validate the quantum computers’ outputs. As quantum computers began to produce results for problems beyond the limits of classical computation, arguments for the validity of the quantum computation always relied on problem instances that were smaller or simpler and extrapolating them to the complexity that classical methods could not reach.
However, these extrapolations do not fully validate the computation in the more complex "advantage" regime—the effects of noise or the propagation of errors can simply be very different.
These three papers represent only part of the ongoing work tracked in the Quantum Advantage Tracker. Today, submissions include candidates from Q-CTRL, BlueQubit, Birla Institute of Technology and Science, Pilani, and we expect the back-and-forth to continue as the community continues to benchmark.
A new way to verify random circuit sampling problems
Random circuit sampling (RCS) has become one of the leading demonstrations of quantum computational separation because carefully constructed random circuits rapidly become intractable for classical simulation. But this hardness also creates a verification problem.
Cross-entropy benchmarking (XEB) requires computing ideal output probabilities, which can become prohibitively expensive for the largest circuits. Earlier experiments therefore often relied on smaller or simplified circuits to infer the performance of circuits that could not be checked directly at full scale. While this provided a practical estimate of hardware performance, it remained a “proxy of a proxy” rather than a direct certification of the classically hard computation itself.
A new paper from researchers at IBM and UChicago addresses this verification gap using a structured alternative called doped Clifford sampling. The authors establish hardness guarantees under assumptions similar to those used for RCS, but crucially, the added structure can also detect errors during the computation.
The Clifford circuit is embedded in a spacetime code, whose detecting regions extend across both the qubits and the circuit’s evolution.

Spacetime codes use ancilla qubits distributed across both space and time to detect errors during a quantum computation. By post-selecting runs that satisfy these consistency checks, researchers can substantially improve the fidelity of logical operations. In this demonstration, a 70-logical-qubit computation achieved roughly a 10× reduction in effective gate error while maintaining practical execution rates.
The authors demonstrated the approach with a large T-doped circuit, in which non-Clifford T gates are strategically added to an otherwise efficiently simulable Clifford circuit. This makes the computation classically difficult while preserving the structure needed for error correction and validation.
The key innovation is that validation becomes part of the computational framework itself. The experiment begins with an encoded Clifford reference circuit, whose output can still be efficiently simulated classically to establish a trusted baseline. The hard computation is then created by strategically introducing T gates while preserving the same syndrome checks.
By combining the fidelity of the classically verifiable Clifford reference with information about syndromes and logical errors, the researchers derive a rigorous lower bound on the fidelity of the encoded logical computation. Rather than relying on an external proxy metric after the experiment, the computation carries enough information to certify its own quality.
This shifts verification from trusting statistical proxies to certifying the fidelity of the logical computation itself, an important step toward fault-tolerant quantum computing, where confidence is established at the logical level.

(L): By introducing T gates into encoded Clifford circuits, researchers created classically hard sampling problems that quickly outpace leading classical simulation methods. (R): Spacetime-code validation remains effective in the hard-computation regime, enabling rigorous fidelity bounds and providing direct evidence that the logical computation remains trustworthy as complexity increases.
Revealing new quantum phenomena that classical methods missed
Meanwhile, researchers at Qedma, working with partners at RIKEN and BlueQubit, studied Floquet dynamics: how an interacting quantum system responds to repeated pulses of energy.
Using circuits of up to 74 qubits, they tracked the system’s magnetization over time and observed persistent oscillations that were not seen in two state-of-the-art classical simulation methods that used one of the world's largest supercomputers at RIKEN. In the most demanding regime, the classical methods disagreed with one another and failed to provide a consistent, reliable answer, while the quantum computation continued to resolve the dynamics.

(L): Error-mitigated quantum computations reveal long-lived Floquet dynamics beyond the reach of leading classical simulation methods. (R): The same behavior is validated using independent error-mitigation approaches and a separate quantum hardware platform, strengthening confidence in the quantum result.
To validate the target circuits without relying on a classical reference, the experiments used Qedma’s QESEM software on IBM Quantum systems with both heuristic and rigorous, unbiased error-mitigation methods. The agreement between these independent estimators, together with consistent results from partially repeating the experiment on a Quantinuum system, provided evidence that the observed behavior reflected the underlying physics.
When you can't verify the answer, verify the process
Researchers at Algorithmiq developed a quantum algorithm to estimate the operator Loschmidt echo, or OLE, a quantity that tracks how information spreads through heterogeneous quantum systems. Their 56-qubit experiments reached regimes where methods from at least three leading classical simulation groups produced inconsistent predictions—and also disagreed with the quantum results.
The central contribution of the work was a strategy towards how the quantum computation could be trusted despite the absence of a reliable classical answer. The researchers applied the same heuristic error-mitigation method across effectively five quantum computers with different noise profiles and found consistent results, making the quantum computation the most credible description among the methods considered.
They then went beyond these consistency checks by showing how rigorous error mitigation can produce unbiased estimates with quantitative error bars when an accurate model of the device noise is available. This reframes the validation problem: rather than asking whether a classical computer can reproduce the quantum result, the task becomes validating the model used to describe and cancel the quantum computer’s errors.

(L): Quantum and leading classical methods produced conflicting predictions for the operator Loschmidt echo, leaving no reliable classical benchmark. (R): The computation was validated through consistency across multiple hardware and noise configurations, together with rigorous error-mitigation techniques that yielded unbiased observables and statistically valid confidence intervals.
The classical benchmarking continues
Announcing quantum advantage does not close the case—it opens the results to a new level of scrutiny.
Last year, a collaboration of classical and quantum computing organizations debuted the Quantum Advantage Tracker, a tool to help the community monitor promising candidates for quantum advantage and systematically evaluate how they stack up against leading classical-only methods.
Since its launch, the tracker has led to fruitful debates that have pushed the limits of both classical and quantum computers. These three experiments live alongside other advantage candidates on the Quantum Advantage Tracker for the broader community to further pressure test.

Today, advantage tracker submissions include candidates from Q-CTRL, BlueQubit, and Birla Institute of Technology and Science, Pilani. We expect the back-and-forth to continue as the community continues to benchmark.
Ultimately, knowing that you trust your computer is essential to establish ground truths. Quantum computers are tools for scientific discovery—and for that to be true, we must trust their outputs. Now, we’re entering the era where quantum is going beyond producing results inaccessible to classical computers—but doing so with trustworthy outputs.




