The AI Industry is Grading Its Own Homework. It’s a Disaster.
We are using LLMs to evaluate other LLMs because humans are too slow. But when two AIs agree, it doesn’t mean they’ve found the truth—it just means they share the same blind spots. Here’s why our automated evaluation pipeline is a systemic risk.