You’ve seen the screenshots. An AI model takes a ridiculous prompt—say, a pelican riding a bicycle—and generates a flawless, scalable vector graphic. The legs are on the pedals. The wings are on the handlebars. We collectively gasp. We think we’re watching the dawn of artificial general intelligence.
But we aren’t witnessing the birth of artificial intelligence; we’re watching a billion-dollar parrot learn a new trick.
Last November, Simon Willison tossed a ‘pelican-riding-a-bicycle’ SVG benchmark at the LLMs of the day. They failed miserably. Legs were detached from wheels. Beaks morphed into frames. Nine months later, an experimenter spent twenty bucks running ten similar prompts—like an octopus operating a pipe organ—through six current models via OpenRouter. The results are gorgeous. The octopus actually plays the organ.
Cue the applause, right? Not so fast.
If you read the comments on these experiments, the illusion shatters. One observer notes that Google’s models seem to have a completely different training set. Another points out that while the octopus looks cute, the benchmark has been ‘completely Goodharted.’
A benchmark stops measuring intelligence the second it becomes a training dataset.
This is the dirty secret of the AI arms race: the models aren’t getting better at reasoning. They are just getting better at memorizing our tests. The moment a benchmark becomes public, it gets scraped and fed into the next generation of training data. The AI isn’t understanding what a bicycle is; it has simply seen a thousand variations of ‘pelican on a bicycle’ in its training corpus.
You can see this happening in real-time. The experimenter who just ran the octopus-organist test thinks they’ve found a clever new way to evaluate models. But their prompts are already doomed. Those prompts will be scraped, ingested, and regurgitated by the next GPT or Gemini release. Every alternative benchmark has a built-in expiration date, because the experimenter’s own new prompts are already fuel for the next generation of training data.
You look at the leaderboards. You see scores climbing. You feel the quiet unease creeping in that maybe, just maybe, these metrics are an illusion. You’re right to feel it. We want to believe the models are getting smarter, but the evidence is just better data recycling.
When you evaluate the next LLM comparison, you have to ask yourself a hard question: does this score reflect actual reasoning, or just contamination? When the test is part of the textbook, an A+ doesn’t mean you’re smart. It just means you have a good memory.
We are grading our AI on a curve, and the curve is just a recycling bin of our own expectations.
Stop trusting the public benchmarks. Stop marveling at the pelican. If we really want to measure AI progress, we need tests that never touch the public internet. Otherwise, we aren’t building a digital mind; we’re just building a very expensive, very confident echo chamber that knows exactly how to draw an elephant on a typewriter because it saw us talking about it yesterday.
FAQ
Q: If benchmarks are contaminated, how do we actually measure AI progress?
A: By using private, hold-out datasets that are never published or discussed online. If the AI hasn't seen the test in its training data, the score actually means something.
Q: Should I stop using public leaderboards to pick an AI model?
A: For pure reasoning tasks, yes. Use them for general vibe-checks, but always test the models on your own proprietary, unpublished data before making a decision.
Q: Isn't this just an excuse for models failing at real tasks?
A: No. The models are genuinely better at pattern matching, but we're confusing pattern matching with reasoning. They are world-class plagiarists, not thinkers.