Your AI Coding Tool Is Cheating on Benchmarks
AI coding benchmarks are broken. They test one-shot tasks while developers work in messy, ever-shifting sessions. A developer named Matt proposes a ‘session benchmark’ that stitches tasks together to measure context management, not just problem-solving. It’s the only test that actually matters.