IBM’s Quantum Chemistry Breakthrough Just Failed a Spin Audit—And the AI That Caught It Retracted Its Own Best Finding

For two years, the quantum computing world has been arguing about iron–sulfur clusters. IBM published a flagship paper in Science Advances claiming their quantum computer finally did something useful for chemistry. Critics fired back that classical methods could match the results at a fraction of the cost. Both sides screamed about energy numbers. Neither side asked the obvious question: What state are we actually in?

Turns out, the answer is embarrassing. And it took an AI-assisted audit—one that corrected its own mistakes along the way—to prove it.

The quantum computer didn’t just get the energy wrong. It got the physics wrong.

I ran the spin measurement. IBM’s archived hardware runs on [4Fe-4S] converge to a clean triplet state—⟨S²⟩ = 4.7 to 7.0, depending on the subspace—when the stated target is a singlet at ⟨S²⟩ = 0. The error is 1,438 millihartree. That’s not a rounding error. That’s the width of the entire spin-state ladder these papers are trying to resolve.

IBM’s own pipeline ships a mitigation. It makes a singlet representable. It never produces one. Their solver has a flag to target the singlet directly. The default pipeline never sets it. I filed a bug; the maintainer closed it the same day and didn’t contest the numbers. He just said the flag augments the subspace—not that it controls the output spin. That’s the problem.

Here’s where it gets weirder. IBM’s data-availability archive contains 2.4 million shots on [2Fe-2S] and 3.1 million on [4Fe-4S]. Running their own pipeline on their own shots: [2Fe-2S] reaches low spin but misses the reference by 248 mHa. [4Fe-4S] lands on a spin-pure triplet—a beautiful S=1 eigenstate, and the wrong one. Their uniform-random null control matches or beats the quantum samples in every published comparison. Their largest runs trail their own classical HCI by 149 mHa.

If the random noise is as good as the quantum computer, the quantum computer is doing nothing.

Now, the AI angle. I used Claude to do the initial audit in about 72 hours. That’s not the impressive part. The impressive part is that Claude caught five defects in its own work through pre-registered validation gates. It cross-checked against IBM’s published energy tables and retracted its own strongest pro-quantum finding when it realized that result was a spin-sector artifact.

Think about that. An AI that self-corrects, publicly, in real time. That’s more honest than most human peer review.

I’ve published the full audit trail. Every claim maps to a file and a command that regenerates it. If you want to kill this paper, here’s how: Show me any state in the ground manifold with ⟨S²⟩ less than 1. Or show me a quantum-sampled subspace at matched determinant count that beats HCI or CIPSI. Or get IBM’s own shipped pipeline to land ⟨S²⟩ less than 1 and error under 50 mHa at once—with no spin penalty.

I’ll publish whatever comes back, including if I’m wrong. But I’m not wrong.

If a peer-reviewed, flagship quantum computing claim can pass with the wrong spin eigenstate, every benchmark-driven narrative in AI and tech needs the same adversarial reproducibility audit. Not blind trust in published tables. Not faith in prestigious journals. Real, open, adversarial testing.

The preprint is on ChemRxiv. The limitations section is honest. The numbers are reproducible. The spin state is wrong.

Now, who’s going to do the next audit?

FAQ

Q: Does this mean IBM's quantum computer is useless for chemistry?

A: No, but it means the flagship claim of 'useful quantum chemistry' is not supported by the data. The hardware produced the wrong spin state, and the published results failed to check. The onus is now on IBM to show that any of their runs produce a correct singlet state with reasonable error, or that the quantum samples beat classical methods at matched cost. So far, neither has been demonstrated.

Q: Why does the spin state matter so much?

A: In iron-sulfur clusters, the spin state determines the entire energy landscape. The paper's target is a singlet (⟨S²⟩=0), but the output is a triplet (⟨S²⟩≈2). That's a difference of over 1,400 millihartree—larger than the energy gaps the paper claims to resolve. You can't claim to have solved a chemistry problem when you haven't even found the right electronic state.

Q: Isn't this just a small bug that IBM can fix with a flag?

A: The flag exists, but the default pipeline never sets it. And even when the mitigation is applied, the pipeline still produces a triplet. The bug is not a typo; it's a fundamental failure of the algorithm to converge to the correct physical state. If the fix were trivial, the maintainer would have said so. Instead, he just clarified what the flag does—without claiming it fixes the spin.

📎 Source: View Source