Stop Testing LLMs with Trivia. Try This Instead.
We’ve been measuring artificial intelligence the wrong way. Static benchmarks just test if a model memorized the internet, but true intelligence requires dynamic reasoning. By forcing LLMs to play Twenty Questions, we finally test if they can ask the right questions, not just recite the right answers.