Stop Trusting LLM-Generated Code. The Security Benchmarks Are a Lie.
We are deploying LLM-generated code at a massive scale, but the security benchmarks we rely on are fundamentally broken. Current tests evaluate isolated snippets, ignoring the reality that security is an emergent property of the entire agentic pipeline. If we don’t start testing how agents scan full codebases, we are flying blind.