You’ve been sold a very specific story about the AI revolution. We are told the battle is won by whoever has the most Nvidia GPUs or the smartest model architectures. But while you were watching the chip wars, a hidden market just exploded—growing 100x in a single year.
AI training data was once the cheap afterthought of model development. Today, it is the most critical bottleneck in the industry, birthing a multi-billion-dollar sector overnight. Companies like Menlo Ventures are tracking 28 data startups collectively pulling in $8.5 billion in revenue, with valuations approaching $100 billion. The sheer scale of growth—Handshake jumping from $5 million to $1.1 billion in annualized revenue in twelve months—creates a visceral sense of urgency. You are witnessing the early days of a new gold rush, and the picks and shovels look nothing like what you’d expect.
Compute is the engine, the model is the chassis, but data is the fuel. Without the fuel, all you have is a ten-million-dollar sports car sitting useless in a garage.
The reason data companies are suddenly printing money is that data’s role has fundamentally shifted. It used to be a one-time raw material. Labs would buy a dataset, train a model, and wait a year before needing more. But model iteration is now breakneck. OpenAI went from taking 13 months between GPT-4 and GPT-4o, to releasing GPT-5.4 and GPT-5.5 just one month apart. Anthropic dropped two flagship models in 11 days.
When models iterate that fast, data purchasing goes from a quarterly project to a daily, high-frequency demand. The labs are constantly discovering capability gaps, and they need highly specific, high-value data to patch them immediately. Scale AI now delivers labeled data via API in under an hour. The old project-based, low-margin commodity business is dead. It has been replaced by a continuous, high-value service.
But here is the twist that everyone is missing: the kind of data that matters is changing.
For years, the industry relied on Type 2 data—tasks designed by humans and labeled by experts. It was controllable and scalable. But as models have gotten smarter, they’ve maxed out what artificial tests can teach them. They don’t need to learn how to take a test anymore; they need to learn how real work gets done.
Enter Type 1 data: raw, unvarnished real-world workflow data.
When the model needs to learn how the real world actually works, artificial data labeling suddenly looks like an expensive game of fill-in-the-blanks.
This is why Handshake—a 12-year-old college recruiting platform—is suddenly a darling of the AI world. They aren’t building models. They are acting as a pipeline, paying thousands of freelancers and enterprises to upload native Word docs, Excel sheets, and code repositories. They control the supply of real-world work, and they are selling that flow directly to AI labs. In a world starved for authentic workflow data, the company that owns the pipes holds all the power.
We are also seeing the rise of Reinforcement Learning (RL) environments. Static datasets are textbooks; RL environments are training grounds where agents explore, fail, and learn to execute tasks. OpenAI is spending tens of millions building simulated website environments. Mercor just acquired Deeptune to build high-fidelity sandboxes of Excel and Salesforce. The era of conversational AI is over; the era of agentic AI is here, and agents can’t be taught with static text—they have to be trained in simulated reality.
Look at the four giants dominating this space—Scale AI, Surge AI, Mercor, and Handshake. They are taking radically different paths, but moving in the exact same direction: up the value chain. Scale is becoming an enterprise AI general contractor. Surge is doing ultra-premium safety evaluation for just five top labs. Mercor is building a global network of PhDs and lawyers. Handshake is dominating the data pipe.
They all realized the same brutal truth. Basic data labeling is a commodity. If you just throw human labor at tagging text, you will be crushed by automation and shrinking margins. The real money—the 50%+ gross margins that Scale AI enjoys—comes from deeply participating in model research, evaluation, and engineering.
Pure-play labeling is dead. The only low-margin link in the AI supply chain will be the one that refuses to move up the value chain.
If you work in tech, invest, or build products, stop obsessing over the models themselves. The models are becoming a commodity. The next wave of value creation is happening in the infrastructure that feeds them. The winners of this AI cycle won’t be the companies writing the best algorithms. They will be the ruthless pragmatists who own the data pipes, control real-world workflows, and understand that in the age of AI, whoever controls the fuel dictates the pace of the race.
FAQ
Q: Why is real-world workflow data (Type 1) suddenly more valuable than expert-labeled data (Type 2)?
A: Because AI models have essentially maxed out what artificial tests can teach them. They already know how to 'take a test.' To evolve from chatbots to autonomous agents, they need to learn how actual work gets done in messy, real-world environments.
Q: What's the practical implication for companies sitting on massive internal datasets?
A: Your internal operational data is no longer a byproduct; it's a highly monetizable asset. Firms like Protege and Handshake are actively paying enterprises to package their internal workflows into AI-ready formats. You are sitting on an untapped revenue stream.
Q: Isn't AI data labeling just a low-skill commodity business that will be automated away?
A: Pure-play, low-skill labeling is absolutely dying. The survivors aren't selling labor; they are selling deep research, evaluation, and engineering services. The data companies printing money right now operate with 50%+ margins because they act as an extension of the AI labs' R&D teams, not as cheap outsourcing.