Stop Trusting AI Leaderboards. They’re Just Benchmaxxing.
AI models are getting terrifyingly good at taking standardized tests, but terrible at solving real problems. We’re trapped in an arms race of ‘benchmaxxing’ where public leaderboards measure overfitting, not intelligence. If you want to know if an AI is actually useful, you have to stop looking at the scores and start looking at the failure modes.