I saw a number that stopped me cold. Doubao’s monthly active users hit 528 million in June — a record. Then I checked QuestMobile’s same-month report. Same product, same month: 382 million. That’s a gap of 146 million users. Enough to swallow DeepSeek’s entire MAU.
Nobody is lying. The numbers are real. But they’re measuring different things. And that’s the problem: the metric we’ve trusted for 15 years — monthly active users — is quietly breaking down on AI products. And most of the industry is still using it to make funding, hiring, and product decisions.
Let me show you three failure modes I found in the latest AI app rankings.
Failure 1: MAU Is Now a Parasite Metric
Doubao lives inside Douyin. Qwen lives inside Taobao and Gaode. Yuanbao lives inside WeChat. This generation of AI apps doesn’t grow independently — they parasitize existing super-apps.
When a user opens Douyin to watch a video and accidentally taps an AI button, is that an AI user? One research firm says yes. Another says no. That’s your 146 million gap.
MAU used to measure product stickiness. Now it measures how much traffic your parent app gives you. It’s a distribution metric, not a quality metric.
Failure 2: Buzz and Retention Have Divorced
Qwen had the highest buzz (social mentions) for three consecutive months. Yet its MAU never beat Doubao. Doubao was quiet — but sticky.
Yuanbao is the brutal example. During Chinese New Year, Tencent threw a massive red-packet campaign. Yuanbao’s DAU hit 40.5 million on New Year’s Eve. By February 23, it was 7.68 million — an 81% crash in one month.
You can buy buzz. You can’t buy retention.
Failure 3: The Analysts Themselves Can’t Keep Their Rulers Straight
I was about to write that Yuanbao’s MAU dropped 40% from 114 million (February peak) to 79 million (June). Sounded powerful. Then I checked the sources. The 114 million came from Tencent’s own internal reporting. The 79 million came from Xsignal. Different instruments.
I’m not hiding this. It’s the perfect example of the problem: even a serious industry analysis can mix two different rulers without realizing it.
Comparing Yuanbao properly with the same source: MAU was 42 million in March, dropped to 32.9 million in September. Then the New Year campaign repeated the same crash. The same hole, stepped into twice.
The Duopoly That Amplifies Everything
Here’s the real kicker. Of the 13 AI apps with MAU > 10 million, ByteDance and Alibaba own 7. Their combined MAU is 10.19 billion — 74.4% of the total top-13 list.
When 74% of the metric comes from two companies’ ecosystem traffic, MAU stops measuring product quality. It measures which super-app is giving you the biggest entrance.
When your product lives inside a daily-800-million-user app, reaching 100 million MAU is not a victory. It’s a side effect.
I’ve Been Here Before, Just Smaller
Reading these numbers, I flashed back to a content moderation project I ran. Offline evaluation: 94% accuracy. Proud number. Deployed to production: dropped 15 points overnight.
I spent days refreshing dashboards, hoping the sample was too small. It wasn’t. The root cause was that our evaluation criteria perfectly matched our labeled dataset — but not real users. We measured a proxy, not the real thing.
The proxy worked in the lab. It broke in the wild. And we were making decisions based on it.
That’s exactly what the AI industry is doing right now, at a billion-user scale. MAU is a proxy. Buzz is a proxy. Token consumption is a proxy. Every shiny metric is a stand-in for value, and the stand-in will eventually break.
The Industry Has Already Changed Its Ruler Three Times in Six Months
First they used DAU. Then OpenAI moved to TPD (tokens per day). Then Gartner said token metrics are misleading — they measure cost, not output. 100k tokens that fail a task are worse than 10k tokens that complete it.
Now the hot new thing is DAA (Daily Active Agents) — how many agents complete a task loop in a real scenario. Google’s CEO proposed it. Gartner backed it.
Three rulers in six months. Each one invalidates the previous. This isn’t progress. This is an industry that hasn’t yet built its own measurement system, grabbing whatever looks solid.
What You Can Do Right Now
Three actions, no theory.
1. Before you trust any AI product number, ask three questions: Who counted it? Which entry points were included? What user behavior does this number actually correspond to? The 146 million gap disappears after the first two questions.
2. In your own product, split your dashboard into two columns. Left column: proxy metrics you report. Right column: the real value you aim for. For each proxy, write down the condition under which it will decouple from real value. For Doubao’s MAU, the decoupling condition is: when embedded entry traffic can’t be attributed. For my content moderation project, it was: when user judgment disagrees with labeling rules.
Teams that can write those decoupling conditions in advance will catch problems six months before teams that just watch the numbers go up.
3. For every key metric, add a counter-metric. MAU up? Check retention curve. Token consumption up? Check task completion rate. Buzz up? Check next-day return rate. Yuanbao’s crash was visible on Day 1 — if anyone had put DAU and retention on the same graph.
After six months of looking at AI app data, my conclusion isn’t about who won or lost. It’s that everyone’s ruler is slightly bent, and we’re making very expensive decisions with these rulers.
Funding rounds. Hiring plans. Product roadmaps. All based on numbers that are true and false at the same time.
My content moderation team eventually rebuilt the evaluation set by sampling real user feedback. Accuracy dropped, but the gap between lab and production never exceeded 5% again. It took three extra weeks.
The industry probably needs the same fix. It just might take more than three weeks.
FAQ
Q: Why does the 146 million MAU gap matter if both numbers are technically correct?
A: Because they measure different things — one includes embedded AI entries in super-apps, the other doesn't. The gap shows that MAU has become a distribution metric, not a product quality metric. Decisions based on either number are decisions based on an incomplete picture.
Q: What should I use instead of MAU to evaluate an AI product?
A: Start with task completion rate. Ask: is this AI actually solving a real problem for users? Then layer on retention, cost per completed task, and user satisfaction. No single metric works — use a balanced set of proxies and explicitly note where each proxy can decouple from real value.
Q: Isn't it normal for metrics to evolve as an industry matures?
A: Yes, but the speed here is alarming — three different rulers in six months, each one invalidating the last. That's not evolution, it's panic. The industry hasn't yet found a stable measurement system, so using any current metric as a definitive signal is dangerous.