You’ve probably seen the headlines: OpenAI’s AI just solved ten advanced math problems for $2000. The implication is clear—machines are now cheaper than mathematicians.
But before you start panicking about your job or your child’s future, let me tell you what OpenAI didn’t say. And why the $2000 figure is the most dangerous number in AI right now.
The $2000 figure isn’t a cost—it’s a marketing number. It’s the kind of number that gets retweeted, screenshotted, and turned into LinkedIn posts about the ‘inevitable obsolescence of human intellect.’ It’s also, almost certainly, a lie.
Here’s the thing: I’ve been watching the AI benchmark game for years. And every time a company drops a headline-grabbing cost-per-solution number, the fine print is where the real story lives. OpenAI didn’t spend $2000 total. They spent $2000 on inference—the final step of running the model. They didn’t include the millions in training costs, the salaries of the mathematicians who curated the problems, the engineering time to set up the pipeline, or the compute for the hundreds of failed attempts that preceded the ten successful ones.
If you bought a car for $2000 but had to spend $200,000 on the factory to build it, you wouldn’t call it a $2000 car. But that’s exactly what OpenAI is doing.
This is benchmark p-value hacking at its finest. Pick the problems that work, report only the successes, and hide the denominator. One commenter on the announcement nailed it: ‘I doubt that they simply pointed their model at these ten specific problems alone and gave the model one shot; therefore the $2000 number could be completely misleading, similar to P-value hacking.’ Exactly.
And the lack of transparency? It’s stunning. OpenAI didn’t release the full experiment logs, the number of trials, or the criteria for problem selection. Without that, we’re not looking at a scientific result—we’re looking at a press release dressed up as research.
Let me be blunt: If you can’t reproduce the experiment, it’s not science—it’s theater.
Now, I’m not saying the math results aren’t impressive. They are. Solving a problem from a top math conference is a genuine achievement. But the way OpenAI is packaging it—as a cheap, inevitable replacement for human mathematicians—is designed to provoke an emotional reaction, not a reasoned discussion.
I’ve seen this movie before. The ‘AI is coming for your job’ narrative is a fundraising tool, not a forecast. The real cost of these breakthroughs is hidden in R&D budgets, compute clusters, and the salaries of PhDs who spend months setting up experiments. The $2000 figure is a sleight of hand.
Does this mean AI won’t eventually transform mathematics? No. It will. But not because it’s cheap. Because it’s getting genuinely better at reasoning. The danger is that we chase the wrong signals—the flashy cost-per-result metrics—and ignore the messy, expensive, non-deterministic reality of how these systems actually work.
So the next time you see a $2000 claim, ask: What’s the real cost? How many tries? What was excluded? Headline achievements often obscure the messy, expensive, and non-deterministic reality of current AI systems.
That’s the truth nobody wants to tweet.
FAQ
Q: What does the $2000 figure actually cover?
A: It covers only the inference cost—the final run of the model on the ten selected problems. It excludes training costs, engineering salaries, problem selection, and the hundreds of failed attempts that preceded the successes.
Q: Is the AI really capable of solving advanced math problems?
A: Yes, the results are genuinely impressive. But without full transparency about the experimental setup, number of trials, and selection criteria, we can't evaluate how robust or generalizable these results are.
Q: Should mathematicians be worried about their jobs?
A: Not yet. AI is becoming a powerful tool for mathematicians, but the 'cheap replacement' narrative is overblown. The real cost of AI research is still enormous, and human expertise remains essential for interpreting results and setting up meaningful problems.