Stop Waiting for AI to Be ‘Good Enough.’ It Never Will Be.

You’ve felt it. That quiet, gnawing anxiety every time a new model drops. GPT-5, Claude 4, Gemini Ultra — the names change, the benchmark charts get prettier, but the question stays the same: Is it good enough yet?

Colin Raffl asked this question on his blog. The top comment, the one that rose above all the thoughtful analysis, was a single word: Never.

That comment isn’t cynical. It’s the most accurate prediction anyone has made about AI. And if you’re building, investing in, or waiting on language models to cross some magical threshold before you act, that one word should terrify you.

Here’s the paradox everyone in AI is living through right now. Language models are simultaneously over-hyped and under-delivering. A model can write a passable sonnet about your dog, summarize a 40-page legal brief in seconds, and code a working React app — then confidently tell you that 2 + 2 equals 5 because it once read a sarcastic Reddit thread. The gap between what these systems can do and what we trust them to do is where entire careers and companies get stuck.

The bottleneck was never the model. It was always us.

Think about what ‘good enough’ actually means. Good enough for what? For drafting an email? Sure, we passed that milestone two years ago. For diagnosing a rare disease? We might be there technically — but no hospital will deploy a system that hallucinates 1% of the time when that 1% means a wrong prescription. The threshold isn’t a number on a benchmark. It’s a moving target shaped by use case, error tolerance, and something nobody wants to admit: human emotion.

We don’t trust AI the way we trust other software. When Excel miscalculates a formula, we blame Microsoft and file a bug report. When an LLM makes something up, we blame the entire concept of AI and question whether any of this works. That’s not a technical problem. That’s a trust problem. And no amount of parameter scaling fixes trust.

I’ve watched teams spend months waiting for ‘the next model’ that would finally make their product viable. They had a prototype that worked 85% of the time. They killed it because 85% wasn’t ‘good enough.’ Six months later, a competitor shipped something that worked 80% of the time — and raised a Series B on it.

The competitor understood something the perfectionists didn’t: perfect is a destination. Useful is a direction. The companies that win pick the one that ships.

This is the twist nobody talks about. The real question was never ‘when will models be good enough?’ The real question is: ‘how do we design systems that work even when models make mistakes?’ The answer isn’t bigger GPUs. It’s guardrails, human-in-the-loop checkpoints, graceful failure modes, and — most importantly — the radical acceptance that your AI will sometimes be wrong.

Every transformative technology was deployed before it was ready. The first cars broke down constantly. Early airplanes crashed. The internet in 1995 was a wasteland of broken links and dial-up screeches. None of that stopped adoption because people understood the trade-offs and built around the flaws.

We’re at that moment with AI. The models we have today are the worst they’ll ever be — but they’re also the best anyone has ever had access to. Waiting for perfection isn’t strategy. It’s procrastination dressed up as prudence.

The teams that figure this out will build the next decade of products. The ones still waiting for GPT-6 to finally be ‘good enough’ will be reading about those products in a newsletter, wondering how they missed it.

You don’t need a perfect model. You need a system that’s honest about its imperfections and useful anyway. That’s the bar. It’s already low enough to step over — if you’re willing to stop waiting and start building.

FAQ

Q: But aren't current models still too unreliable for serious use cases?

A: For some tasks, yes. But 'unreliable' and 'useless' aren't the same thing. The question isn't whether models fail — it's whether your system can handle failure gracefully. Add guardrails, human review, and clear failure modes. Ship the 85% solution before someone else ships the 80% one.

Q: What does this mean for my AI product roadmap?

A: Stop benchmark-watching and start user-testing. Your customers don't care about MMLU scores. They care whether your product saves them 30 minutes a day. Design for the mistakes your model makes today, not the perfection you hope it reaches tomorrow.

Q: Isn't shipping imperfect AI irresponsible?

A: Shipping AI without guardrails is irresponsible. Shipping nothing while competitors eat your market is also irresponsible. The responsible move is building systems that are transparent about their limitations — not waiting for a model that never lies. Perfectionism isn't ethics; it's procrastination with a moral alibi.

📎 Source: View Source