You’ve seen the headlines. AI models are passing the bar exam, writing flawless code, and generating poetry that would make Shakespeare weep. We are constantly told the singularity is just around the corner. But put these trillion-parameter brains in a cockpit, and they suddenly lose their minds.
Over on Twitch, a live experiment pitted two of the smartest AI models on the planet—GPT-5.6 Sol and Kimi K3—against Kerbal Space Program. The goal? Speedrun a basic orbital flight. The reality? The smartest machines on Earth spent hours floating aimlessly in the vacuum of space, completely baffled by gravity.
We laugh at this. One Twitch commenter perfectly summed up the vibe: “They’re both just stuck in space LOL.” But underneath the schadenfreude is a glaring, uncomfortable reality. We are building brains that can talk, but they have absolutely no concept of how the physical world works.
We treat language tests as the ultimate benchmark for intelligence. If an AI can write a Python script or ace a medical exam, it must be smart, right? Wrong. Language is a map of reality, but current AI architectures are hopelessly lost in the map, completely blind to the territory. They know the word “orbit,” but they don’t feel the physics of it.
A human child playing with a ball understands trajectory, mass, and gravity before they ever learn the word “physics.” Biological brains are embodied. We learn through scraped knees, dropped toys, and falling off bicycles. AI doesn’t have a body. It has a text corpus. It learns the statistical likelihood of the word “fall” appearing next to “gravity,” but it has no intrinsic spatial model to apply that concept in real-time.
This isn’t a bug that a few more billion parameters will fix. You cannot scale your way out of a fundamental architectural blind spot. GPT-5.6 and Kimi K3 failing at a decade-old video game isn’t a temporary setback; it’s the glass ceiling of the current paradigm.
So the next time you read a headline about an AI achieving human-level performance on a standardized test, remember them spinning endlessly in a digital void. Until an AI can land a rocket, don’t let anyone tell you it understands the real world.
FAQ
Q: Isn't Kerbal Space Program notoriously hard? Why is this a fair test?
A: Yes, KSP is hard for a human learning it from scratch. But an AI with access to the exact physics equations and orbital mechanics should calculate it instantly. It didn't fail at math; it failed at spatial reasoning.
Q: What does this mean practically for AI development?
A: It means we need to stop obsessing over language benchmarks and start focusing on embodied AI—models that interact with simulated physics, not just text tokens. If your AI can't apply logic to 3D space, it's not ready for the real world.
Q: Are you saying Large Language Models are a dead end?
A: Not a dead end, but a cul-de-sac. They are incredible linguistic engines, but pretending they are on the verge of general intelligence just because they talk well is a massive, dangerous delusion.