Someone Just Built a Master Key for Every AI Agent. The Response Is Terrifying.

You probably haven’t heard of pleasedonotescape.com. You should have. It’s a filterable, comprehensive list of AI agent jailbreaks — every known method for making an AI do what it was explicitly designed not to do, cataloged, tagged, and served up like a menu.

And the tech community’s response? “Cool! Looking forward to checking it out.”

That’s it. That’s the collective reaction to what might be the most consequential dual-use artifact in AI safety today.

A list of jailbreaks isn’t a warning label. It’s a blueprint. And pretending otherwise is how we sleepwalk into the exact catastrophe everyone claims to be worried about.

Here’s the thing that should make you uncomfortable: this list exists in a strange liminal space. On one hand, it’s exactly what AI safety researchers have been begging for — transparency, shared knowledge, a common vocabulary for vulnerabilities. On the other hand, it’s a ready-to-use offensive toolkit for anyone with a grudge, a curiosity, or a profit motive.

The Hacker News thread that birthed this project started innocuously enough. People were trading names of AI agent sandboxes and jails. Someone — we don’t know who — decided to compile them all. Filter them. Organize them. Make them searchable.

That act of compilation is more powerful than people realize.

When you catalog every known way to break out of a jail, you’re not just documenting the jail — you’re defining its walls. You’re telling the entire ecosystem: here’s what counts as confinement, and here’s what counts as escape.

Think about that for a second. The person who builds this list is implicitly shaping the norms of acceptable AI behavior. They’re drawing the map that everyone else will navigate by. Researchers will use it to patch holes. Attackers will use it to find new ones. Regulators will use it to argue for restrictions. And every one of them will treat the list’s boundaries as the boundaries of the problem.

But what about the jailbreaks that didn’t make the list? What about the ones that are still sitting in private Discord servers, or in the notes of a red team that got acquired before publishing? The list is comprehensive — until it isn’t. And the confidence it inspires might be the most dangerous thing about it.

I’ve watched this pattern before in cybersecurity. Every time someone publishes a vulnerability database, two things happen simultaneously: defenders get better, and attackers get organized. The net effect depends entirely on who moves faster. In traditional infosec, defenders usually have a head start because patches ship before disclosures go public.

AI doesn’t work that way. A jailbreak technique, once known, can be replicated across thousands of agents in minutes. There’s no patch Tuesday for a language model that’s already deployed in production. The asymmetry favors the attacker to a degree that makes traditional CVE disclosure look quaint.

In software security, the bug exists in the code. In AI security, the bug exists in the conversation — and you can’t patch a conversation.

So where does that leave us? The list at pleasedonotescape.com is genuinely valuable. If you’re building AI agents, deploying them, or just using them, you should absolutely study it. You should understand the attack surface. You should harden your systems.

But you should also understand what it represents: a moment where the AI community decided that transparency was worth the risk, without fully reckoning with what that risk looks like when the technology in question can be weaponized at the speed of text.

The real question isn’t whether this list should exist. It does. The question is whether we’re mature enough as an industry to use it responsibly — or whether “Cool! Looking forward to checking it out” is the best we can muster when someone hands us the keys to every lock.

The most dangerous jailbreak isn’t the one in the list. It’s the complacency that comes from believing the list is complete.

FAQ

Q: Isn't publishing jailbreaks just responsible disclosure, like in traditional cybersecurity?

A: No. In traditional infosec, you can patch the code before or simultaneously with disclosure. AI agents can't be patched the same way — a jailbreak technique replicates at the speed of text across already-deployed models. The attacker-defender asymmetry is fundamentally different.

Q: Should I actually use this list if I'm building AI agents?

A: Yes, absolutely. Study every entry. Understand your attack surface. But treat it as a starting point, not a complete map. The jailbreaks that aren't on the list are the ones that will actually hurt you.

Q: Is the real problem here that someone built this list, or that the community doesn't take it seriously enough?

A: Both, but the latter is worse. Building the list is an act of transparency with real trade-offs. Responding to it with 'Cool, looking forward to checking it out' reveals an industry that hasn't internalized that its own safety mechanisms are made of glass.

📎 Source: View Source