You’ve probably felt it. That uneasy whisper when yet another AI agent benchmark hits a new high score. Is this real progress, or are we just getting better at measuring the wrong thing?
The answer is uncomfortable: most of what we call ‘model improvement’ is actually just harness improvement.
I spent the weekend inside a new paper called Computer Anthology, and it didn’t just present a benchmark. It dissected the very act of benchmarking. And what I found shook how I think about every AI agent leaderboard I’ve ever trusted.
Let’s start with the raw data. The researchers compared two evaluation harnesses—let’s call them Terminus and Codex—running the same model (GPT-5.5). The result? Swapping harnesses gave a performance boost equivalent to upgrading an entire model generation. GPT-5.5 with Codex scored roughly the same as GPT-5.6 with Terminus—at a lower cost. That’s not a tweak. That’s a confession.
We’ve been treating benchmarks as if they’re neutral, objective rulers. They’re not. They’re measurement infrastructure as opinionated as the models they test. The harness determines how the agent sees the environment, what actions are allowed, how success is scored. Change the harness, and you change the game.
This is where the tension really bites. Benchmarks need to evolve to stay relevant—otherwise agents just memorize the test. But evolution destroys comparability. The moment you update a benchmark, you can’t compare scores to last year. The field is caught in a paradox: we need progress, but we also need consistency. You can’t have both.
I watched a commenter on the paper nail it: “Refreshing to see something practical instead of another leaderboard battle.” Yes, because this paper exposes the battle behind the leaderboard. The real fight isn’t between models; it’s between measurement frameworks. The next time you see a jaw-dropping AI agent score, ask yourself: What harness was used? Would the same model score 20% lower with a different one?
Here’s the twist you didn’t see coming: this isn’t a critique of AI progress. It’s the most hopeful thing I’ve read in months. Because if we can improve performance by simply fixing the harness, that means we’re nowhere near the ceiling. The low-hanging fruit isn’t model architecture—it’s measurement design. The field is partly measuring its own infrastructure, and that’s okay. It’s how science works.
But it also means every vendor claiming a ‘model generation leap’ needs to show you the harness. Otherwise, you’re buying a benchmark score, not a capability.
So here’s my take: Neutrality is dead. You either evaluate your agents with a clear understanding of the harness effect, or you’re being misled. The Computer Anthology team didn’t just give us a benchmark. They gave us a mirror. And in that mirror, we see that the biggest breakthroughs in AI agents might not come from the models at all. They’ll come from the invisible infrastructure that decides what counts as success.
FAQ
Q: What question would a skeptic ask?
A: Doesn't this just mean we need better benchmarks? Yes, but better benchmarks aren't neutral either. The real insight is that every benchmark is a lens, and lenses introduce distortion. The gold standard is reporting results with multiple harnesses and being transparent about the measurement infrastructure.
Q: What's the practical implication?
A: If you're buying or building on AI agents, the harness choice can change results as much as the model choice. Don't compare scores from different benchmarks or different harness versions. Demand full transparency: which harness, which version, which task definitions. Otherwise you're comparing apples to orchestrated fruit.
Q: What's the contrarian take?
A: The harness effect is actually great news. It means we have a powerful lever for improvement that doesn't require expensive model training. The field should double down on building better measurement infrastructure—it's cheaper, faster, and more accessible than chasing bigger models. The real AI progress might come from the evaluators, not the evaluated.