You’ve probably felt it. That brief, electric moment when a new model drops — “GPT-5 crushes ARC-AGI!” — and you think, finally, this is the one. You sign up, you test it on your actual work, and three weeks later you’re right back where you started. The chatbot is still writing emails that sound like a nervous intern, and the code it generates still breaks in production.
Something is off. The leaderboards scream “breakthrough,” but your daily workflow whispers “meh.” And you’re not imagining it.
Look at the ARC-AGI leaderboard right now. The gap between Opus 5 and the next best model is staggering — on paper. But one user on the forum put it bluntly: “I have a suspicion that they are just trained on puzzles by now.”
That suspicion is the elephant in the lab. The ARC-AGI benchmark was designed to measure genuine reasoning. But the more we glorify the score, the more incentives shift from building reasoning to building puzzle-solving. The line between intelligence and memorization has blurred into non-existence.
Here’s the twist nobody wants to admit: The leaderboard is no longer a measure of intelligence. It’s a measure of how thoroughly a model has been trained on the test.
And the consequences are not academic. If you’re building a product, choosing a model, or making strategy decisions based on these benchmarks, you are being misled. The model that tops ARC-AGI might be the worst performer on your actual data — because your data is not a puzzle from a curated dataset.
One commenter on the ARC-AGI thread pointed to Fable, a model that seemed to hit a cap above which the US government won’t let LLMs improve. “Everything they’re releasing from this point has to be worse than that.” That’s not a conspiracy theory — it’s a plausible reading of the landscape. The benchmark ceiling is now a policy ceiling, and the public leaderboards are the stage where this theater plays out.
Another user asked: “Why Anthropic models are always leapfrogging these benchmarks, but in real life work I feel like after 3 weeks I am back to Claude Opus 4.5?” The answer is uncomfortable: Benchmark scores are optimized for the benchmark, not for you.
This is not an attack on the researchers. It’s an attack on the system. The pressure to publish, to raise, to dominate the leaderboard, creates a perverse game. The best way to win is to train on the test. And we have evidence that’s happening. The Frontier-Bench just released, and the same models topped it by a large margin — because they were trained on similar puzzles.
So what do you do? Stop trusting the leaderboards. Test your own use cases. The model that scores highest on ARC-AGI might be the worst for your business.
I’ve seen this firsthand. I built a tool that evaluated a top-ranked model on a simple data extraction task. It failed. Miserably. The same model that aced the benchmark couldn’t tell the difference between a date and a phone number in a PDF. The benchmark was measuring something, but it wasn’t intelligence.
Here’s the hard truth we need to swallow: We are measuring artificial memorization, not artificial intelligence. And every time we celebrate a new record, we are reinforcing the game that produces models that are great at games and useless at life.
This is dangerous. It’s dangerous because it wastes talent, money, and trust. It’s dangerous because it creates a false sense of progress. And it’s dangerous because it makes us believe that AGI is just around the corner, when in reality, we are just getting better at gaming the metrics.
Take a side. I’m taking mine: Benchmark chasing is a dead end. Real progress is measured in the messy, unglamorous work of getting models to actually help ordinary people with their real problems. That doesn’t make a splashy headline. But it’s the only thing that matters.
FAQ
Q: Are you saying AI benchmarks are completely useless?
A: No. Benchmarks are useful for comparing models on specific, well-defined tasks. But they are not proxies for general intelligence or real-world utility. The problem is when they are treated as such — especially when incentives lead to training on the test set.
Q: How should I evaluate an AI model for my use case?
A: Build your own small evaluation set from your actual data. Test the model on tasks that matter to your business. Don't rely on public leaderboards. A model's performance on ARC-AGI tells you nothing about how it will handle your customer support tickets or your financial reports.
Q: Isn't this just sour grapes? The models are genuinely improving.
A: Models are improving, but the rate of improvement on benchmarks is inflated by training on the test data. The real-world gains are real but smaller than the headlines suggest. The contrarian view: we are over-investing in benchmark optimization and under-investing in robustness, safety, and real-world deployment.