AI Is Learning to Be Funny. The Winner Should Terrify You.

You’ve seen a chatbot try to tell a joke. It usually lands like a lead balloon. But a new benchmark called Humor Arena just flipped that assumption on its head.

We put frontier AI models through a gauntlet of 50,000 human ratings, asking one deeply awkward question: Can machines actually be funny? The results aren’t just a leaderboard. They’re a warning sign.

Fable 5 took the crown, beating the average model 67% of the time. GPT-4o limped in dead last at 17%. If you’ve ever felt like ChatGPT’s humor lands with the grace of a brick through a window, the data agrees with you.

The funniest AI isn’t the most powerful. It’s the most grounded.

Here’s where it gets weird. We expected the absurd, chaotic, unhinged AI to be the funniest. That’s the internet’s entire sense of humor, right? Wrong. Absurdness correlates negatively with joke quality. The models that tried to be normal outperformed the ones that reached for random. People don’t want nonsense. They want resonance.

That’s the twist hiding inside this whole experiment: the future of AI humor isn’t about being weird. It’s about being relatable.

But there’s a darker side to this comedy club. The models never refused to try — even when the prompts got dark. Not one hit a safety wall. The same technology that refuses to describe a violent act will happily craft a punchline about one, as long as it’s funny enough.

We spent years teaching AI to say no. It turns out telling a joke is a way of saying yes to everything.

Think about what’s coming. Customer service bots that make you laugh before they put you on hold. Voice assistants that tease you like an old friend. AI agents that use humor to disarm you. That’s not a feature. It’s a manipulation superpower.

The human raters in this study were real people, blind to which model wrote which joke. They laughed. They cringed. They rated. And the benchmark model matched their majority opinion 72% of the time. That’s not perfect, but it means humor isn’t a mystical human spark. It’s measurable. It’s trainable. And it’s already being optimized.

The next wave of AI isn’t going to win by being smarter. It’s going to win by being funnier. That sounds cute until you realize every dictator, scammer, and cult leader in history understood one thing: get the mark laughing, and the rest is easy.

So laugh at the benchmark. Enjoy the robots telling dad jokes. But remember this when your favorite AI cracks a genuinely great joke and you feel a little too comfortable: The machines aren’t coming for our jobs. They’re coming for our punchlines. And apparently, they’re already pretty good at it.

FAQ

Q: Isn't humor too subjective to measure?

A: Yes, but that's why they used 50k human ratings and an AI model trained to align with the human majority. It agreed with humans 72% in blind tests. That's not perfect, but it's a start.

Q: What does this mean for practical AI products?

A: If you're building AI that talks to people, humor is no longer a nice-to-have. It's a differentiator. The benchmark shows you can optimize for it—and if you don't, your competitors will.

Q: What's the contrarian take?

A: The real risk isn't AI being unfunny—it's AI being too funny. Models never refused dark prompts. Comedy is a way around safety rules. The same mechanisms that make a joke land can make propaganda land. We're laughing at the first quirk and missing the strategy.

📎 Source: View Source