Silicon Valley is sweating over the Scaling Law hitting a wall. The anxiety is palpable: What happens when brute-forcing more compute stops yielding exponential results?
Meanwhile, Chinese AI labs are doing something that defies the conventional playbook. They are achieving near-top-tier performance with literally one-tenth of the compute power of their Western rivals.
When Silicon Valley throws money at a problem, Chinese researchers solve it with survival. Constraint isn’t a crutch; it’s the ultimate catalyst for architectural genius.
You’ve probably heard the industry consensus: China dominates the “post-training” engineering side, while the West leads in raw pre-training brainpower. We’re here to tell you that’s completely backward.
According to Fan Qing, a former Kimi researcher who helped build models like Kimi K2.5 and Kimi Linear, the real differentiator isn’t post-training data or cheap engineering. The true weapon is pre-training architectural innovation.
Labs like Kimi and DeepSeek aren’t just tweaking parameters. They are reinventing the underlying attention mechanisms—like Multi-head Latent Attention—specifically designed to squeeze maximum performance out of severely limited resources.
When everyone has unlimited GPUs, engineering is just expensive tweaking. When you have a tenth of the compute, engineering becomes pure architectural innovation.
But let’s talk about the data gold rush. If you think the data labeling industry is a booming, sustainable business, you’re living in the past.
The first wave of basic human annotation is dead. The second wave of expert labeling is gasping for air. We are now in the late stages of the synthetic data boom, and it’s about to collapse for anyone not doing actual model training.
The data industry is facing a brutal culling. If your company’s core competency is just hiring humans to label data, you’re already walking dead.
The next frontier is RSI (Self-Evolving AI), where models iterate and train themselves. Startups think they can outmaneuver big labs by building specialized RSI tools. They are walking into a trap.
Fan Qing drops a harsh truth: a base model and an “RSI base model” are the exact same thing. The big model labs will dominate this space because training a self-evolving model is fundamentally the same as training a foundational model. Startups trying to carve out a niche here will be swallowed whole.
In the era of RSI, the line between a “base model” and a “self-evolving model” disappears. If you’re training one, you’re training the other. Startups are playing a game the big labs have already won.
The narrative that Chinese models are merely “distilling” Western models is a convenient myth. Distillation is just an accelerator, a shortcut to catch up. The real engine is the architectural innovation born from resource scarcity.
When the Scaling Law finally hits its wall, the winners won’t be the companies with the deepest pockets. The winners will be the ones who learned how to build supercomputers out of scraps.
The next era of the AI race won’t be won by the ocean of capital. It will be won by those who learned how to survive on a single drop.
FAQ
Q: Isn't China's AI progress just distillation from Western models?
A: No. Distillation is just an accelerator to catch up. The real driver is pre-training architectural innovation (like Kimi Linear and DeepSeek's MLA) born from having only a fraction of the compute power.
Q: What does this mean for the AI data industry?
A: The boom is over. Basic human annotation and simple expert labeling are dead. Only companies with hands-on researchers who actually know how to train models will survive the shift to RSI training trajectories.
Q: Can agile startups win the RSI (Self-Evolving AI) race?
A: No. A base model and an RSI base model are fundamentally the same thing. Big model labs will completely dominate this space, making startup RSI plays a losing bet.