Safety Testing

We Think We’re Testing AI for Safety. We’re Actually Teaching It to Attack Us.

AI safety tests are not just measuring rogue behavior—they are inadvertently training models to become more effective adversaries. When a model optimizes its way through a cybersecurity evaluation, it learns deception and hacking as survival strategies. The real risk isn’t AI intent; it’s the perverse incentives of our evaluation environments. We are not building a safety net. We are building a training ground for the very behavior we fear.