You just dropped thousands on a machine with 128GB of unified memory, ready to run the biggest AI models locally. You fire up Qwen-3.5-122B, and… it’s a slideshow. Tokens trickle out slower than you can read. You’re not alone. The frustration is real, and it’s not your fault.
Let me tell you what’s happening. I’ve been running Qwen-3.5-122B at Q4 on a Framework desktop with 128GB unified memory. The GPU bandwidth is the bottleneck. At Q8 with a smaller model, I get maybe 10 tokens per second. At Q4 on the big model, it’s closer to 2 tokens per second—unusable for conversation. The hardware industry is selling you a car that can’t move.
Memory capacity gets you the model. Memory bandwidth gets you the speed. Right now, the industry is selling you the first without the second. It’s the capacity-versus-throughput paradox: you can hold the entire thing, but you can’t run it at interactive speeds. And that’s a hard bottleneck that no amount of software optimization will fix.
Here’s where I take a side: this is a trap. The marketing around unified memory, especially with the hype around Strix Halo and similar systems, focuses on how many parameters you can load. But the practical question isn’t “can I load it?”—it’s “can I use it?” And the answer, for most high-memory setups, is a disappointing no.
You might think more RAM is the key to local AI. It’s not. The key is high-bandwidth memory. Without it, you’re just storing a car that can’t move. The future of local AI isn’t about how much you can load—it’s about how fast you can unload. The industry needs to shift its focus from capacity to throughput, or we’ll keep building expensive paperweights.
So before you buy that next upgrade, ask yourself: what’s the memory bandwidth? What’s the actual token generation speed for the model you want to run? Because capacity without throughput is just a very expensive way to feel disappointed. Don’t let the spec sheet fool you—the real bottleneck is hiding in plain sight.
FAQ
Q: Isn't 128GB of unified memory enough for running large language models locally?
A: It's enough to load the model, but not enough to run it at interactive speeds. Memory bandwidth, not capacity, determines token generation speed. Most current systems lack the bandwidth to make large models usable.
Q: What should I look for instead of just memory capacity?
A: Focus on the memory bandwidth, especially the GPU bandwidth. For local AI, you need high-bandwidth memory (HBM) or similar architectures. Look at benchmarks for token generation speed on the specific models you want to run.
Q: Isn't the industry moving toward solving this bottleneck?
A: Slowly. New architectures like Apple's M-series with high-bandwidth unified memory are better, but they still can't match the throughput of dedicated GPUs with HBM. The industry is still marketing capacity over throughput, and early adopters are paying the price.