You’ve spent weeks agonizing over which AI model to use for your coding agent. Claude? GPT-4o? DeepSeek? You stare at the benchmark leaderboards, make your choice, and pray to the silicon gods that your dev velocity will finally double.
But you’ve been duped. Here’s the dirty secret of the AI coding boom: you aren’t benchmarking intelligence. You’re benchmarking the scaffolding.
Think about it like cars. If Car A is faster than Car B, you might assume Car A has a superior engine. But it could just have better tires, a smoother gearbox, or superior aerodynamics. In the world of AI coding agents, the engine is the LLM. The tires, gearbox, and suspension? That’s the harness—the tools, the planning logic, the execution environment.
A recent empirical study on harness design just blew the lid off this illusion. The researchers found that the exact same harness feature that rescues a weaker model can be completely unnecessary—or actively harmful—for a stronger one. It’s a moving target. Planning improves success rates for dumb models at a higher cost, but for smart models, it just adds overhead and reduces success rates.
Giving a frontier model a rigid step-by-step planner isn’t like giving a master chef a recipe. It’s like giving a Formula 1 driver a backseat driver. It just slows them down.
We saw this firsthand when Anthropic quietly disabled their Todo/task-tracking tools for their newer models. Why? Because the model was smart enough to track its own state. Forcing it to use external scaffolding was an insult to its intelligence. Weaker models need the notepad; strong models just get confused by the bureaucratic overhead.
Then there’s the ‘bash-capable’ myth. We talk about models being ‘bash-capable’ as if it’s a personality trait. It’s not. It’s an emergent interaction between the model and the environment. You can’t separate the model’s capability from the harness it’s running in. They are a single, entangled system. Predefined tools help models with weak bash control, but if a model already knows what it’s doing, bash-only yields higher success at a lower cost.
Most AI leaderboards aren’t measuring who is the smartest. They’re measuring who wears the best-fitting suit.
And the suits they’re wearing in these benchmarks are often outdated. The current research is lagging the frontier. Studies are still running on obsolete tool defaults, and many completely exclude the actual frontier models like Qwen and DeepSeek. You’re making multi-million dollar architectural decisions based on data from models nobody uses anymore, tested in environments that no longer matter.
For anyone building or buying AI coding agents, this is both terrifying and deeply reassuring. It’s terrifying because you can’t just swap in a new LLM and expect magic—the same harness will break differently across models. But it’s reassuring because it means your agent’s success isn’t just a function of OpenAI or Anthropic’s latest training run. It’s a function of your engineering.
Your harness choices will make or break the model. Stop shopping for the smartest engine. Start building a better chassis.
FAQ
Q: What is the key takeaway?
A: See the article.