You’ve seen the hype cycle a thousand times. A new AI paper drops, promises to revolutionize deep learning, and then immediately face-plants on a basic benchmark. We roll our eyes, mutter that backpropagation is forever, and move on to the next parameter-bloated model.
But what if we are actively ignoring the exact mathematical blueprint of the human brain just because it can’t beat a 1998 baseline right out of the gate?
Enter Sakana AI’s latest research on Augmented Lagrangian Predictive Coding. The top comment on the release perfectly captures the industry’s dismissive stance: “~85% accuracy on MNIST. Sigh. How does it do on ImageNet?”
The community is already writing its obituary. 85% on MNIST is objectively terrible. A basic convolutional neural network from a decade ago crushes that number with barely any effort. But in our rush to mock the underwhelming empirical performance, we are missing the theoretical earthquake happening right under our feet.
Backpropagation didn’t conquer the world because it was elegant; it conquered the world because we threw millions of dollars of GPUs at it until it stopped sucking.
We treat backprop as the undisputed king of machine learning, forgetting that it sat in the academic doghouse for decades because it couldn’t scale on the hardware of the 1980s and 1990s. Now, we have a biologically plausible alternative—an alternative that aligns with Karl Friston’s Free Energy Principle—and we’re ready to discard it because it doesn’t instantly scale to ImageNet. It’s the equivalent of throwing away the Wright brothers’ flyer because it didn’t break the sound barrier on its first flight.
What makes this Sakana paper so thrilling isn’t the 85% accuracy; it’s the theoretical ambition. This isn’t just another gradient descent variant. It’s a mathematically rigorous attempt to explain how the brain actually solves the credit assignment problem without a magical, biologically impossible backward pass.
We are treating biological plausibility like a bug, when it’s the only feature that might actually get us to AGI.
The AI community suffers from a severe case of empirical myopia. If a new architecture doesn’t immediately shatter state-of-the-art records on massive datasets, we dismiss it as a toy. But backpropagation is a biological impossibility. Your brain does not have a mechanism to propagate error gradients backward through trillions of synapses in perfect continuous time. It just doesn’t happen. If we want sample-efficient, robust, and truly intelligent systems, we have to start looking at how the only general intelligence we know of—us—actually learns.
If your model learns exactly like a human brain but can’t classify a cat in 4K resolution yet, the AI community calls it a failure. The brain calls it Tuesday.
Predictive coding via augmented Lagrangian methods is rudimentary. It’s clunky. It achieves 85% on MNIST. But it points toward a unified theory of intelligence that bridges artificial intelligence and neuroscience. We can either keep brute-forcing backpropagation with exponentially more compute until the power grid collapses, or we can fund and nurture the biologically plausible paradigms that might actually unlock the next leap. The choice is ours, but the clock is ticking.
FAQ
Q: 85% on MNIST is objectively terrible. Why should anyone care?
A: Because backprop was also terrible until we had the compute to make it work. The point isn't the score today; it's that this algorithm is biologically plausible and aligns with how the brain actually functions.
Q: What does this mean for current AI development?
A: Nothing today. You won't be swapping out your PyTorch backprop loops tomorrow. But in 5-10 years, this framework could be the foundation of AI that learns sample-efficiently without requiring a power plant to train.
Q: Is backpropagation actually a dead end?
A: For achieving true AGI on a reasonable energy budget? Yes. It's a brute-force hack that scales only with exponentially more compute. Biology has a better way, and we should probably start paying attention to it.