Model Compression Is a Distraction. Rethink Your Loss Instead.
You don’t need a $30,000 datacenter GPU to distill AI models. The real bottleneck isn’t your model size—it’s your loss function. By chunking and fusing KL-divergence passes, we can drop VRAM usage from O(n²) to O(n), democratizing AI research for anyone with a 6GB consumer GPU.