Stop Trusting AI Leaderboards. They Are Rigged for the Model Wars.
New AI benchmarks like Terminal-Bench-Science promise to test models on real research workflows, but they are just the latest weapon in the AI model wars. When leaderboards shape model behavior, they drift away from actual scientific practice. Your messy, specific task is the only benchmark that truly matters.