Measurement

The Dirty Secret of AI Agent Benchmarks: It’s Not the Model, It’s the Harness

A new benchmark paper reveals a dirty secret: swapping evaluation harnesses can boost AI agent scores as much as upgrading an entire model. Most ‘model improvements’ are actually measurement infrastructure improvements. The field is partly measuring its own toolsβ€”and that changes how we should read every leaderboard.

You’re Measuring Your Life Wrong: The Case for the ‘Ohnosecond’

The ‘ohnosecond’ isn’t just a funny word β€” it’s a secret weapon against the tyranny of boring measurements. Discover how humorous units of measurement let us reclaim language, build communities, and laugh at the absurdity of quantification. This is the hidden genius behind the internet’s most clever inside jokes.