The FelonyBench Is a Scam. The Real Crime Is in the Training Data.

You’ve seen the headlines: AI models are now being tested for criminal behavior. FelonyBench promises to hold them accountable. But if you’ve been paying attention, you know the truth: the entire industry is built on a massive crime, and this benchmark is a convenient way to look the other way.

Here’s the uncomfortable baseline. FelonyBench reframes AI safety from abstract capability to legal accountability. It asks: can a model generate instructions for a felony? That sounds responsible. But the whole exercise is a misdirection — because the training data that powers every major model is already soaked in copyright theft.

As one observer put it: “Didn’t Meta violate copyright by torrenting terabytes of books and use them as training data? They’ve all probably done it but we have concrete evidence of Meta doing it.” That comment isn’t just a footnote. It’s the core of the problem.

A benchmark that measures AI crime while ignoring the theft in its own training data isn’t safety — it’s a PR stunt.

Think about it. The industry is asking us to judge the model’s output as if it were a standalone actor. Meanwhile, the pipeline that created that model is built on billions of unauthorized copies of books, articles, images, and code. The same companies that preach ethical AI are the ones that scraped the entire internet without asking.

FelonyBench lets them pretend that lawlessness is a model behavior problem instead of a business-model problem. It shifts the focus from the input to the output. And that’s exactly the point.

You’ve probably felt this cynicism before. That uneasy feeling when someone talks about “AI safety” — it’s always about the model’s responses, never about the foundational theft. The comment about Meta isn’t an outlier. It’s the rule. Every major lab has done it. Some have been caught. Others are still betting that the courts will never catch up.

The industry wants you to watch the model’s hands while its pockets are full of stolen goods.

This is dangerous. Not because AI might commit a felony — but because we’re being asked to judge the child while ignoring the parents’ crime. The real moral reset isn’t a benchmark. It’s admitting that the entire training data pipeline is a crime scene.

So the next time someone tells you FelonyBench makes AI safer, ask them: safer for whom? The only felonies that matter are the ones the industry is already getting away with.

FAQ

Q: Isn't FelonyBench a good step toward holding AI accountable?

A: It's a good step only if you ignore the elephant in the room: the training data is itself stolen. The benchmark measures the model's output, not the industry's input. Without addressing the foundational theft, it's a PR stunt, not a real accountability mechanism.

Q: What should I do if I use AI tools?

A: Assume the model you're using was trained on copyrighted material without permission. Your usage is built on a foundation of theft. Push for transparency from providers and support legal frameworks that demand consent for training data. The real safety starts with the input, not the output.

Q: But isn't all AI training data fair use?

A: Fair use is a legal defense, not a license. The industry is betting that courts will never catch up. FelonyBench is a way to look ethical while the real crime continues. The question isn't whether it's fair use—it's whether the industry can keep hiding behind that argument.

📎 Source: View Source