You can’t strip-mine the collective knowledge of humanity, bake it into a proprietary server, and then cry foul when someone else extracts value from your exhaust.
Yet, here we are. Y Combinator president Garry Tan just kicked the hornet’s nest by demanding that US open-weight AI labs be allowed to “distill” frontier models. Distillation is essentially a smarter, cheaper way of learning from a massive AI model to create a smaller, highly capable one. It’s how the open-source community keeps up with the billionaires.
And the billionaires are furious.
Proprietary labs like Anthropic and OpenAI want to lock down their models. They want to control who can learn from them, how they’re used, and—most importantly—who gets to pay the toll. They argue that allowing open-weight labs to distill their outputs is unfair, maybe even illegal.
But let’s get one thing straight: The proprietary AI labs didn’t ask permission when they vacuumed up the entire internet to train their models in the first place.
They scraped copyrighted books. They ingested personal blogs, forums, and code repositories without a dime in compensation. They built their trillion-dollar valuations on the silent labor of millions of people who never signed a terms of service agreement.
Now that they have the castle, they want to raise the drawbridge.
But the copyright framing is a distraction. It’s a shiny object designed to make you debate intellectual property while the real heist happens in the background. The real battleground isn’t about who owns the data. It’s about who controls the infrastructure.
We’ve seen this movie before. During the dot-com boom, telecom companies laid thousands of miles of fiber optic cable in public rights-of-way. They used public land, public subsidies, and public spectrum to build their networks. Then, they tried to monopolize the bandwidth.
What happened? The Telecom Act forced them to lease that fiber to competitors at a fair price. Because when you build essential infrastructure, you don’t get to gatekeep it.
If a model is smart enough to replace human labor, it is no longer just a piece of software—it is essential infrastructure. And infrastructure doesn’t belong in a walled garden.
Anthropic and OpenAI want you to believe they are protecting the integrity of their systems. They want to claim the moral high ground, arguing that orderly distillation protects safety. But you can’t claim moral superiority over a model you built on stolen data.
What they are actually protecting is their moat. They are terrified of a world where a scrappy startup can distill GPT-4 or Claude 3, run it locally on cheap hardware, and undercut their API pricing by 90%.
The debate over AI distillation isn’t about protecting artists; it’s about protecting toll booths.
If the proprietary labs win this fight, the AI economy doesn’t just slow down—it consolidates. Every startup, every indie developer, every enterprise will be forced to pay a tax to two or three gatekeepers. Innovation dies not because of a lack of ideas, but because the cost of compute becomes a protection racket.
We should be doing everything in our power to encourage distillation. We need an ecosystem of smaller, efficient, open-weight models that anyone can audit, modify, and build upon. Preventing token-consumers from developing competing products isn’t just bad for business—it should be litigated as anti-competitive behavior.
The internet was built on the commons. AI was built on the commons. It’s time we treat frontier models like the telecom lines they are: shared substrates that belong to the economy, not to the CEOs who happened to wire them up first.
FAQ
Q: Isn't it unfair for open-source labs to just copy the work of proprietary labs?
A: No, because the proprietary labs built their work by scraping the entire internet without permission or payment. You can't claim theft when your entire foundation is built on stolen property.
Q: What happens if proprietary labs successfully ban distillation?
A: The AI economy consolidates around two or three gatekeepers. Startups will be forced to pay massive API tolls, and innovation will be bottlenecked by corporate monopolies controlling the rails.
Q: How is AI like telecom infrastructure?
A: Just like telecom companies laid fiber in public rights-of-way and were forced to share it, AI labs used public data to build essential infrastructure. Frontier models should be treated as common carriers, not private property.