You’ve done it. We all have. You write a bold claim, hesitate, and paste it into ChatGPT or Claude to ‘fact-check’ it. The AI gives you a confident green light. You hit publish. But what if I told you that green light is completely arbitrary?
New research from Lenz, published on Zenodo, just exposed a massive, quiet crack in our AI workflows. They tested frontier LLMs to see if they could act as interchangeable arbiters of truth. The result? They can’t. In fact, they violently disagree on factual claims, and there is absolutely no independent ground truth to settle the score.
We built a generation of machines that can argue brilliantly, and then we asked them to be our judges.
We’ve been obsessing over the wrong threat. Everyone is terrified of hallucinations—when an AI confidently makes up a fake statistic or a nonexistent legal case. That’s the obvious danger. You can usually catch a hallucination if you look closely. The real, insidious danger is what happens when the models agree.
If you ask three different AIs to verify a complex claim, and they all say ‘True,’ you feel safe. But Agreement is not evidence of truth; it is often just evidence of shared training data.
When these frontier models—whether it’s GPT-4, Claude, or newer entrants like Fable and Sol—agree on a fact, they aren’t tapping into some universal fountain of truth. They are often just echoing the same blind spots baked into their overlapping datasets. You aren’t fact-checking; you are echo-chambering.
And when they disagree? It’s even worse. You’re left standing there with two highly confident machines giving you completely different realities, and no independent way to prove which machine is right. We outsourced our sense of reality, and the machines returned the favor with a shrug.
The real danger isn’t the hallucination you catch. It’s the blind spot you inherit.
If you are using AI for research, writing, or decision-making, you need to understand this immediately. LLM-based fact-checking is not neutral. It is a highly biased lens that inherits the flaws of the specific model you chose. Treating any single model as the ultimate arbiter of truth introduces a massive, unacknowledged epistemic risk into your work.
Stop treating AI like a magic oracle. It’s just a very persuasive debater. And right now, the debaters are arguing amongst themselves, hoping you won’t notice.
FAQ
Q: If AIs disagree, doesn't that just mean we need a better AI?
A: No. A 'better' AI still operates on training data, not universal truth. You're just trading one set of blind spots for a more confident set of blind spots.
Q: So how do I actually verify facts now?
A: Use AIs for synthesis and brainstorming, not arbitration. If a fact matters, verify it through primary sources. The AI is a starting point, never the final word.
Q: Isn't human fact-checking just as flawed?
A: Yes, but humans know they are flawed. AI presents its flaws with absolute, unearned mathematical confidence, making its failures far more dangerous.