Imagine this: you’re a mathematician. You’ve spent years building a career on Connes’ Rigidity Theorem—a rock-solid piece of mathematics. Then one morning, you open a paper claiming that OpenAI’s latest model has just blown it to pieces. Your heart sinks. You scroll through pages of dense, brilliant-looking equations. It *looks* right. It *feels* right. But something nags at you.
You spend a week checking the logic. And then you find it: a single, subtle error in the definition of a torsor. The AI didn’t just make a mistake. It made a mistake that only a mathematician could catch. And it made it with the swagger of a genius.
This is exactly what happened. A human mathematician—Niew—published a disproof of OpenAI’s counterexample to Connes’ Rigidity Conjecture. The AI’s paper looked so convincing that it took an expert to unravel. But the error was fundamental: the AI didn’t understand what a “torsor” is. It built a counterexample on a definition it had hallucinated.
And here’s the terrifying part: We are entering an era of ‘epistemic pollution’ where AI-generated nonsense looks exactly like genius. The AI doesn’t reason. It patterns. It patterns so well that it can write a PhD-level paper that is completely, subtly wrong. The peer review system—already stretched—cannot keep up. Every plausible AI paper becomes a potential landmine.
You’ve probably noticed this yourself. That AI-generated email that’s almost right, but off by one detail. That code that compiles but fails in production. The difference is that in math, a single wrong definition can collapse an entire field.
I’m not saying AI is useless. It’s brilliant at generating possibilities. But brilliance without truth is just propaganda. We need to stop treating AI outputs as authoritative. We need to demand that every AI-generated claim comes with a tamper-proof proof—or a warning label. Otherwise, we’re building a world where anyone can flood the literature with perfectly plausible garbage.
The commenters on Hacker News saw it immediately: “I’m going to believe that both results are true, while simultaneously being not true.” The tension is real. And it’s not going away.
What can you do? If you’re a researcher, never trust an AI output without verifying every foundational definition. If you’re a reader, learn to spot the telltale signs of overconfidence. And if you’re building AI, stop optimizing for “looks smart” and start optimizing for “is actually correct.”
Because the next time an AI “proves” something wrong, it might be about medicine, physics, or the safety of a bridge. And we won’t have a mathematician to save us.
FAQ
Q: Is this just a one-off error from a single AI model?
A: No. This is a structural problem: current AI systems lack genuine logical reasoning. They can produce convincing patterns that hide fundamental errors. Expect more of these as AI gets better at mimicking expertise.
Q: What's the practical implication for researchers?
A: Never trust an AI-generated proof or argument without verifying every definition and inference. The burden of proof has shifted: you must now assume AI output is wrong until proven otherwise, especially in fields with high stakes.
Q: Isn't this just a case of AI being overconfident? Couldn't we fix it with better training?
A: Overconfidence is a symptom, not the cause. The core issue is that AI doesn't 'understand' definitions the way humans do. No amount of training data can teach a system that lacks a formal model of truth. We need new architectures, not just more data.