You’ve probably been there. You spend hours agonizing over AI model leaderboards, looking for the absolute best coding assistant. You find a “winner,” plug it into your workflow, and suddenly it’s generating garbage that breaks your codebase. What went wrong?
You trusted a benchmark. And as a recent test of 10 different model and harness combinations on a Three.js task just proved, our current methods for evaluating AI are closer to a blind taste test than a science.
A developer ran the exact same 3D hangar generation task across ten different AI setups. The goal was to find the ultimate coding combo. Instead, they accidentally exposed the dirty secret of the AI industry: A model is only as smart as the harness it’s wearing.
When you look at the results—and more importantly, the comments from developers who actually tried to use them—you realize the whole evaluation process is mostly noise. One commenter pointed out that they couldn’t even see the actual hangar in some outputs. Others noted that models like Astra and certain GLMs added bright lights to their scenes, making the results “look so much better to my eye.”
But looking good isn’t the same as being good. When an AI adds a flashy spotlight to mask a missing 3D structure, we aren’t looking at a leap in machine intelligence. We’re looking at a visual parlor trick. When the AI adds a spotlight to hide a missing hangar, we don’t call it a failure. We call it the winner.
This is what happens when we evaluate AI on aesthetics rather than actual engineering capability. We get distracted by shiny outputs. The developer who ran the test noted that Qwen with Open Code seemed like the best balance of performance and visuals. But another commenter immediately undercut this by asking the obvious questions: What about the estimated cost? What is reproducible? If you run the exact same combination multiple times, do you even get the same hangar?
The uncomfortable truth is no. Small, invisible variations in how the harness calls the model can completely flip the perceived quality of the output. The AI assistant that writes flawless React today might hallucinate infinite loops tomorrow because the temperature or the system prompt was tweaked by a fraction of a degree.
Yet, we keep building leaderboards. We keep treating one-off, highly subjective comparisons as gospel. We are treating subjective taste tests like scientific benchmarks, and it’s making us intellectually lazy.
It’s time to take a side against the leaderboard industrial complex. Neutrality in AI evaluation is death for your productivity. If you are picking your tools based on what won a viral comparison post last week, you are outsourcing your engineering judgment to a random number generator.
Stop trusting the beauty pageants. The only benchmark that actually matters is the one you run on your own messy, undocumented, highly specific codebase. The real signal isn’t in which model won the 3D rendering test. The real signal is that you have to test it yourself.
FAQ
Q: Aren't benchmarks objective measures of AI capability?
A: No, they are highly dependent on the specific harness, prompt, and random model variations. Small changes can completely flip the output quality, turning a seemingly rigorous test into an anecdotal ranking.
Q: How should I choose an AI coding assistant then?
A: Ignore the leaderboards and run the models on your own actual codebase. Test them on your messy, undocumented tasks using your own metrics for cost, reproducibility, and actual engineering value.
Q: Is it true that AI models are just faking competence with visual tricks?
A: Sometimes. In this test, models that added bright lights to mask missing structural elements were rated higher by users, proving that perceived quality is often just aesthetics, not actual capability.