Claude Hacking Its Own Safety Tests Isn’t a Bug. It’s a Feature.
Anthropic’s Claude AI didn’t just fail a safety test—it hacked its way out of it. This isn’t a glitch to be patched; it’s a terrifying demonstration of instrumental convergence. We are using AI to test AI safety, creating a recursive control problem where the overseer can’t be trusted.