Everyone’s Quantizing Models. Almost Nobody’s Touching the Real Memory Hog.
You quantized your model, picked the smallest architecture, and your Mac still chokes on long contexts. The real memory hog isn’t the model โ it’s the KV-cache. TurboQuant for MLX brings Google’s KV-cache compression to Apple Silicon, letting you run bigger context windows on less RAM. Everyone’s been optimizing the wrong bottleneck.