You’ve watched the GPU prices skyrocket. You’ve felt the sting of being locked out of AI research because you don’t have a $30,000 H100 cluster sitting in your basement. We’ve been sold a massive lie: that to train a model from scratch, to actually teach it and align it to your values, you need corporate-scale compute.
But what if the bottleneck isn’t your graphics card at all?
The AI industry is obsessed with fitting the entire brain into a single chip. That’s not how intelligence works—that’s how brute force works.
Enter a frustrated developer who recently posted a project on Hacker News under a name so audacious it borders on mocking the industry: Mini-AGI. He apologized for the pretentious name, but he shouldn’t have. Because the project runs on a pathetic 8GB of VRAM—less memory than your average gaming PC—and it completely breaks the rules of how we’re supposed to train neural networks.
Here’s the dirty secret of mainstream deep learning: it relies on massive, randomized batches of data. To process these batches, you need huge amounts of VRAM to store the weights, the data, and the gradients simultaneously. It’s a hardware arms race designed to keep hobbyists on the sidelines.
So, this developer threw out the batch sizes entirely. He built a Mixture-of-Experts (MoE) model that trains in a continuous, single stream—batch size 1—reading 32,000 characters at a time. Just like you or I reading a book.
By discarding randomized batches, you don’t need a supercomputer. You just need a hard drive.
The experts in this model aren’t static. They are dynamically added and pruned on the fly, loaded and unloaded from your disk as needed. The VRAM is no longer the ceiling. The disk space is. You aren’t constrained by the memory on your GPU; you’re constrained by the storage on your PC.
We’ve been treating VRAM as the size of the brain. We should have been treating it as the size of our working memory, and the hard drive as the library.
Why does this matter? Because alignment is too important to leave to Sam Altman or Sundar Pichai. When you train a model on a corporate cluster, you are aligning it to corporate interests. When you train it on your own machine, reading the data you choose, in the order you choose, you own the alignment process.
The comments on the original post say it all. One user wrote that seeing ‘Mini-AGI’ and ‘8GB VRAM’ in the same sentence is a breath of fresh air. Another just threw a crumpled paper ball at the audacity of it.
It’s audacious, but it’s a glimpse of a future where the gatekeepers lose their keys. The model is still cooking through its first 7.8 billion characters, but the scaling laws look promising. You can clone the repo, run it yourself, and watch it learn in real-time.
The future of AI won’t be built in a data center. It will be built in a garage, running on a graphics card you bought off eBay.
FAQ
Q: Is this actually AGI?
A: No, it's a continuous learning model that can run on 8GB VRAM. The name is half-joke, half-aspiration, but the underlying architecture is a very real, functional approach to local training.
Q: What's the practical implication?
A: You can train and align models locally without a data-center budget. By swapping experts to disk and using batch-1 training, your hard drive capacity becomes the only limit on model size.
Q: What's the contrarian take?
A: Mainstream deep learning's reliance on massive random batches is a massive waste of hardware. Single-stream, batch-1 training is the only way consumer hardware will ever compete with corporate compute.