Alibaba Banned Claude. It’s the Best Thing That Could Happen to Its AI.

You know that feeling when your company forces you to use an internal tool that is objectively worse than the industry standard? The frustration. The clunky UI. The sheer waste of time.

That’s exactly what happened at Alibaba this week. The tech giant officially banned all Anthropic products, including Claude, forcing its entire workforce to switch to its own AI, Qwen. Employees are understandably furious. They went from the silky smooth experience of Claude to a homegrown model overnight.

Most people are reading this as a geopolitics or data compliance play. They’re missing the real strategy. This isn’t about security. It’s about turning 100,000 employees into the world’s most demanding beta testers.

In the AI industry, there’s a bizarre phenomenon: models that never lose a benchmark test but always lose the user experience. Teams optimize for MMLU scores, post a victory tweet, and then watch users silently uninstall the app a week later.

Benchmarks are for egos. Complaints are for products.

When you push a product to external users, the feedback loop is broken. You get vague complaints like “it feels off” or “it’s not as good as X.” But when your own employees are forced to use your AI for 8 hours a day to write code and draft reports, the feedback is instant, brutal, and hyper-specific.

Look at DeepSeek. They didn’t just build a good model; they were backed by a quantitative hedge fund whose researchers used the AI daily for high-stakes trading analysis. When a model hallucination costs millions, you fix it fast. The internal team was the harshest critic, driving rapid iteration that left Silicon Valley scrambling.

Tencent did the same thing. Their TAPD project management tool was used internally for 12 years, battered by thousands of engineers, before it was ever released to the public. It launched as a fully mature product.

If your own team won’t use your product, you’re just running a very expensive science experiment.

Before this ban, Alibaba was actually reimbursing employees for using Claude and GPT. It was a pragmatic approach, but it meant the people building Qwen weren’t even using it. The feedback loop was dead.

Now, the loop is violently reconnected. Engineers will scream that the code completion is dumb. Marketers will complain that the long-context handling drops data. And every single one of those complaints is a golden product requirement document.

The short-term pain is real. But the long-term gain is a feedback loop no external user study can ever match.

For product leaders, the lesson is clear. Stop chasing vanity metrics. Stop relying on low-frequency user surveys. Don’t just ask your team to test your product. Force them to live in it, hate it, and fix it. That’s how you build a moat.

FAQ

Q: Isn't forcing employees to use inferior tools just hurting productivity?

A: Yes, in the short term. But that productivity hit is the price of admission for a high-density feedback loop that will make the product superior in the long run.

Q: What's the practical implication for product managers?

A: You need to migrate your team's actual daily workflows to your own product. Symbolic 'testing' doesn't work; it has to be mandatory, high-stakes usage.

Q: Is dogfooding really the only way to build good AI?

A: No, but it's the fastest way to bridge the gap between theoretical benchmarks and actual user experience. Without it, you're just guessing what users want.

📎 Source: View Source