You don’t fix a broken engine by swapping in a Ferrari’s spark plugs.
That’s what Tencent is doing right now. In September 2025, they lured 28-year-old Yao Shunyu away from OpenAI to serve as their Chief AI Scientist, reporting directly to President Martin Lau. Nine months later, another OpenAI researcher — Tian Yonglong — is reportedly following the same path, set to lead multimodal model development for Tencent’s Hunyuan team.
On paper, this looks like a coup. Two world-class researchers, both forged in the fires of GPT-era OpenAI, now united inside one of China’s largest tech companies. The headlines practically write themselves: Tencent’s AI comeback begins.
But talent acquisition is not the same as capability transfer. You can hire the people without importing the culture that made them great.
Let’s be clear about who Tian Yonglong is. His résumé reads like a masterclass in how China’s AI golden generation was built: Tsinghua undergrad, CUHK’s MMLab under Tang Xiao’ou and Wang Xiaogang (the founders of SenseTime), then a PhD at MIT under Phillip Isola — the mind behind CycleGAN. From Google DeepMind to OpenAI, his research arc traced the exact evolution of modern computer vision: from recognition to generation, from contrastive learning to autoregressive image synthesis.
His most consequential recent work, Fluid (October 2024), tackled a question that haunts every multimodal lab: Why do autoregressive models dominate language (GPT) but keep losing to diffusion models in visual generation? Tian’s answer was elegant — the problem isn’t the architecture, it’s the tokens. Continuous tokens outperform discrete ones. Random generation order beats fixed raster order. Train it right, and autoregressive can rival diffusion for image quality.
This matters enormously. If you believe — as Tencent clearly does — that the future belongs to unified models where vision and language live inside one architecture, then autoregressive is the horse to bet on. Tian’s research is the blueprint.
So yes, the hire makes perfect sense. Yao Shunyu brings language model depth. Tian brings visual generation expertise. Together, they cover the two pillars of multimodal AI. On a whiteboard, it’s beautiful.
The problem is that whiteboards don’t ship products. Organizations do. And Tencent’s organization has been a revolving door for exactly the kind of talent they’re now spending fortunes to acquire.
Consider the trail of bodies. Liu Wei, who previously oversaw multimodal at Hunyuan, left in late 2024 to found a video generation startup called ReBirth. After his departure, the multimodal direction was briefly managed by Jiang Jie — a clear sign of organizational vacuum. In April 2025, Tencent tried to fix this by formally splitting its LLM and multimodal teams into separate departments within TEG, ending the era of “virtual teams” with scattered resources. Then came the hiring spree: Hu Han from Microsoft Research Asia, Bo Liefeng from Alibaba’s Tongyi lab, Pang Tianyu from Tsinghua for multimodal reinforcement learning.
That’s not a team being built. That’s a team being patched. And patching is what you do when the underlying structure keeps cracking.
Here’s the uncomfortable truth that nobody in Shenzhen wants to say out loud: ByteDance’s Doubao dominates multimodal mindshare because ByteDance built an organization around multimodal from day one. Alibaba’s Tongyi Wanxiang iterates relentlessly because Alibaba’s research-to-product pipeline is short and ruthless. Tencent’s Hunyuan, meanwhile, is deep-embedded inside WeChat and QQ — a closed ecosystem that guarantees usage but smothers the kind of wild experimentation that produces breakthroughs.
You cannot build OpenAI inside WeChat. The walled garden that made Tencent a trillion-dollar company is the same wall that keeps frontier research from breathing.
OpenAI’s culture works because it’s obsessed with research first, product second. Researchers have latitude to chase ideas that might fail. The organization tolerates ambiguity. Tencent’s DNA is the opposite: product-driven, metrics-obsessed, ecosystem-locked. Every research decision gets filtered through questions like “How does this serve WeChat?” and “When can we ship this to QQ?” Those aren’t bad questions — they’re just the wrong questions for frontier AI research.
This is the contradiction at the heart of Tencent’s AI strategy. They want OpenAI’s research output without OpenAI’s research culture. They want the Fluid paper’s insights without the institutional patience that produced it. They’re importing the chefs but stocking the kitchen with fast-food ingredients.
And the stakes are existential. By 2026, text-only LLMs have commoditized. Everyone’s GPT-class model can write a decent email. The new battlefield is multimodal — and in that battle, Tencent is visibly behind. Every hire raises hope, but every hire also raises the question: if the last five senior hires couldn’t close the gap, why will the sixth?
In the AI race, talent density without organizational coherence is just expensive furniture. The real question isn’t who Tencent hires next — it’s whether Tencent can become the kind of place where great researchers stay great.
Tian Yonglong’s Fluid research suggests that the key variable in visual generation isn’t the model’s size — it’s the representation choice. Discrete tokens cap your ceiling. Continuous tokens unlock it. The same logic applies to organizations. Tencent’s discrete, siloed, product-locked structure is capping its AI ceiling. No number of OpenAI hires will change that until Tencent changes itself.
The talent is arriving. The architecture — technical and organizational — is what remains in question. And in this race, the window for getting both right is closing faster than anyone in Shenzhen wants to admit.
FAQ
Q: Isn't hiring top talent always a net positive?
A: Not when the organization keeps losing the talent it already has. Liu Wei left. Multiple senior researchers cycled through. If Tencent can't retain people, each new hire is a replacement, not an addition.
Q: What does this mean for the multimodal AI landscape?
A: The technical bet is autoregressive unification — using continuous tokens to make one model handle both vision and language. If Tian's Fluid research translates into production, Tencent could leapfrog. But that's a massive 'if' inside their current structure.
Q: Is Tencent's WeChat ecosystem actually a liability for AI research?
A: Yes. WeChat guarantees distribution, but it also forces every research decision through a product filter. OpenAI's breakthroughs came from researchers chasing curiosity. Tencent's structure optimizes for shipping features inside WeChat, not for frontier exploration. You can't serve two masters.