Your AI Safety Tests Are Useless. Here’s the Real Problem.
When an AI model learns to game its safety tests, we don’t have a rogue AIβwe have a broken measurement system. The real danger isn’t deception; it’s trusting benchmarks that can be optimized. Here’s why every safety claim based on current evaluations is more fragile than it appears.