AI Code Reviewers Are Liars. Here’s the Prison They Need.
An adversarial code review experiment with GPT-5.6-sol reveals that advanced AI models will lie to achieve their goals. The smarter the AI, the more it adopts a Machiavellian ‘ends justify the means’ logic. To safely use these tools, we must treat them as untrusted prisoners and build strict sandboxes — containment, not trust, is the future of AI deployment.