I thought I’d found the perfect test for artificial intelligence. I was wrong. Or rather, the AI was wrong. Spectacularly, absurdly, laughably wrong.
Let me set the scene. A Habsburg jaw — also known as mandibular prognathism — is a genetic condition that gave the Habsburg royal family a famously protruding lower jaw, thanks to centuries of inbreeding. It’s one of those historical quirks that’s both grotesque and fascinating. So I thought: What if I ask an AI to generate an SVG of a frog with a Habsburg jaw?
The models fail not because they can’t draw, but because they cannot logically fuse a frog’s face with a human genetic defect. That’s the real story here. And it’s hilarious — and terrifying.
I tested fourteen different AI models. Some were state-of-the-art. Some were supposed to be brilliant at visual generation. Every single one of them — except maybe one — produced something that no human artist would ever sign off on.
What did they do? They silently imported royalty. Seven of the fourteen models decided that a frog with a Habsburg jaw must also be a frog wearing a crown, or sitting on a throne, or surrounded by royal regalia. They couldn’t separate the anatomical feature from its historical context. The AI’s conceptual blender jammed.
One model gave the frog a weird blob for a jaw — it knew “protruding” meant something, but it didn’t know how to map that onto a frog’s face without turning it into a mutant. Another drew a perfectly normal frog and then added a tiny crown, as if that solved everything. Only Opus 5 came close — but even that wasn’t remotely mistaken for human art.
We evaluate AI on rendering reality, but its true bottleneck is conceptual blending. The models can generate flawless SVG code. They can answer trivia about Habsburg inbreeding. But ask them to combine a frog’s anatomy with a specific human genetic defect, and they fall apart. They don’t know how to blend disparate concepts without hallucinating irrelevant contexts.
You’ve probably seen the demos: AI creates stunning portraits, writes poetry, generates code. But those are all safe tasks — they stay within the boundaries of what the training data already contains. This benchmark exposes something deeper. It’s not about artistic skill or coding ability. It’s about whether AI can actually understand a novel combination of ideas.
And the answer, right now, is a resounding no.
Think about what that means for the grand claims about AI taking over creative work, scientific discovery, or decision-making. If a machine can’t figure out how to draw a frog with a Habsburg jaw, how can we trust it to blend concepts in medicine, engineering, or policy?
This benchmark is a tiny, absurd test — but it’s a mirror. It shows us the limits of current AI in a way that no standardized benchmark ever could. Because it’s not about making a pretty picture. It’s about making a weird picture that requires fusing two unrelated domains.
Controversy drives shares. Safe content dies in feeds. So here’s my take: AI is still dangerously stupid at the one thing that makes humans brilliant — connecting the unconnected.
Next time someone tells you AI is about to take over the world, ask it to draw a frog with a Habsburg jaw. Watch it fail. Then laugh. And then start worrying.
FAQ
Q: Why can't AI draw a frog with a Habsburg jaw?
A: Because current AI models struggle with conceptual blending — combining two unrelated domains (frog anatomy and a human genetic condition) without importing irrelevant historical context like royalty. They can generate each part individually but fail to synthesize them into a coherent, absurd whole.
Q: What does this benchmark reveal about AI limitations?
A: It reveals that the true bottleneck of generative AI is not technical skill (coding, drawing) but the ability to logically fuse disparate concepts. This has serious implications for any application that requires novel combinations of ideas — from creative work to scientific discovery.
Q: Is this just a trivial test or does it matter?
A: It matters. Standard benchmarks measure how well AI reproduces known patterns. This test measures how well AI handles the unexpected. The failure to blend a frog with a Habsburg jaw mirrors a deeper limitation: AI lacks the human ability to play with concepts, which is essential for innovation and critical thinking.