Stop Obsessing Over Token Speed. The Real Local AI Bottleneck Is Apple Silicon’s Memory Bandwidth.
The real bottleneck in local AI on Apple Silicon isn’t token speed—it’s memory bandwidth and software instability. Hardware benchmarks promise 52 tok/s, but real-world usage reveals crashes, OOMs, and broken drafting. Until inference frameworks mature, local AI remains a hobbyist’s playground, not a production tool.