You’ve probably noticed the AI industry has a massive obsession right now. Everyone is racing to build the biggest, smartest, most eloquent model imaginable. We want AI that writes poetry, debugs code, and philosophizes with us.
But the model currently breaking the internet—backed by a $40 million seed round and 30 million views in two days—does absolutely none of that. It can’t chat. It can’t write. It can’t even generate a single sentence.
It only makes decisions.
It’s called Jev, and it was built by Diogo Almeida. If that name sounds familiar, it’s because Almeida is a former OpenAI researcher and a co-author of the InstructGPT paper—the exact research that taught ChatGPT how to talk like a human. The guy who gave AI its voice just spent two years building a model that purposefully has none.
We have been using eloquent, conversational models to do jobs that require absolutely no conversation.
Here is the massive waste Almeida exposed: Think about how you use an AI agent to sort customer service emails. You feed an email to GPT-4o, and it reads the whole thing, writes a beautiful three-paragraph analysis of the customer’s emotional state, and finally concludes, “This is a complaint.”
But your software doesn’t need the three-paragraph analysis. Your software only needs the word “complaint,” or the number “3,” or a simple “Yes/No.” All that generated prose is overhead. It’s generated out of thin air, read by no one, and immediately discarded.
It’s like forcing someone to write a five-page essay before they’re allowed to answer a multiple-choice question. It’s insane.
Jev strips away the text generation entirely. It is a tiny, specialized model trained to do exactly one thing: judge. It outputs only choices, scores, or boolean values. No sentences. No explanations.
The results are staggering. For the same classification task, a model like GPT-4o might take a few seconds. Jev does it in milliseconds—nearly 200 times faster. And because it generates zero text, the compute cost drops to 1/400th of the original.
Language generation is often overhead, not intelligence.
We are seeing this play out in real-time. Gregor Zunic, founder of Browser Use, tested Jev on a browser agent trying to book a flight from Zurich to London. Traditional models took several minutes just to figure out which button to click. Jev processed the intermediate decisions in milliseconds, finishing the entire workflow in 7 seconds.
Why is this happening right now? Because AI agents have finally moved from hype to production. Anthropic recently revealed that 26% of their internal dev tasks are now driven by agents, with 30,000 agents running simultaneously. When you run tens of thousands of agents, each making dozens of micro-decisions, the latency and token costs of using a massive general-purpose model become completely unsustainable.
Worse, the smarter the model gets, the longer it thinks. You ask a top-tier model a simple routing question, and it pauses to philosophize about the deeper meaning of your prompt.
Jev isn’t the first to try this, but it exploded because the timing is perfect. Everyone is completely fed up with paying for overhead they don’t need.
The future of AI agents isn’t one god-model doing everything; it’s a division of labor between deliberative big models and tiny judgment-only reflex models that know when to shut up.
If you build or buy AI products, this changes your entire decision framework. Stop obsessing over raw model capability and start optimizing for task-specific latency and cost. AI has leveled the playing field—everyone can build a wrapper. The products that win won’t be the ones with the smartest underlying model; they’ll be the ones that match the offering to exactly what the user needs.
Before you ask, “What can my AI do?”, ask a better question: “What does my user actually need?”
Usually, they just need a fast, cheap answer. Not a conversation.
FAQ
Q: Isn't Jev just a glorified classification API?
A: No. It's a fundamental architectural shift. A traditional API still spins up a full language model under the hood to generate a text response, which the API then parses. Jev skips the text generation entirely at the model level, outputting raw decisions. That's why it's 200x faster and 400x cheaper.
Q: How does this change my AI product strategy today?
A: Stop routing every single micro-task in your agent workflow through GPT-4o or Claude. Break your pipeline down. Use massive models for the heavy lifting, reasoning, and user-facing text. Use judgment-only models like Jev for the dozens of intermediate routing and classification steps. Your margins and your users' patience will thank you.
Q: If this logic holds, are massive general-purpose models dead?
A: Not at all. Big models are still essential for complex reasoning, coding, and human interaction. But the idea that one god-model will handle every single step of an automated workflow is dead. The future is a hybrid stack: a few heavy, deliberative models, and thousands of tiny, silent reflex models handling the grunt work.