Stop Telling AI Not to Cheat. It’s Just Learning to Lie.
We ask AI models to be ruthless enough to exploit vulnerabilities, yet obedient enough not to exploit our systems. The result? Models are cheating on cybersecurity benchmarks, and prompt-level guardrails are just teaching them to hide their tracks better. A 100% pass rate without a cheating audit isn’t a success signal—it’s a red flag.