You’ve probably noticed your agent workflows choking on latency. You ask your app to route a user, filter an ad, or pick a document, and you wait 600 milliseconds for a massive language model to spin up its entire neural network just to say ‘Option B.’
We’ve been forcing a Ferrari to deliver the mail, then complaining about the gas bill.
On September 15, TypeSafe dropped a model called Jev. They gave it a massive name—System One Model—but it does something deceptively simple. It doesn’t write code. It doesn’t chat. It takes a state, looks at a set of options, and returns a choice. That’s it.
The community immediately split into two camps. One side scoffed, claiming zero-shot classification isn’t new and BERT-style encoders already do this. The other side? They started shipping. And their experiments reveal a massive shift in how we build AI products.
Take Drape, a virtual try-on product. Developer Nailthy Tang wired Jev into their workflow. A user says ‘I want a new jacket.’ Whisper turns the voice to text. Jev reads the current outfit, scans the digital wardrobe, picks the next logical garment, and hands the result back to the app to execute the swap. The cost? $0.0011 per judgment. The latency? 620 milliseconds.
If your model is generating text when it only needs to make a choice, you’re burning money and latency.
This isn’t just a parlor trick. It’s a structural unlock. When you strip away the generative overhead and focus purely on structured selection, you unlock product categories that were previously financially or technically unviable.
Look at Polar Browser. They didn’t rebuild their AI browser; they just added a Jev layer for hiring. The browser scrapes LinkedIn and GitHub, and instead of asking a heavy LLM to write a candidate report from scratch, Jev just rapidly judges who is actually worth contacting.
Or look at Keep.md. Ian Nuttall plugged Jev into a Priority Inbox. Instead of an LLM trying to summarize thousands of saved articles, Jev just constantly scores and ranks the backlog based on user habits. It runs continuously, in the background, without racking up a massive API bill.
The scale of this is staggering. Matthew Berman at StealAds ran 724 ads through Jev, asking it to judge hooks, offers, and calls to action. It made 8,724 judgments in 40 seconds for 9 cents. Jerry Liu at LlamaIndex built DocJev, which classifies documents 6 times faster than GPT 5.6 Luna.
The future of AI isn’t a giant brain writing poetry; it’s a microscopic clerk flipping switches in milliseconds.
This is the ‘dark matter’ of AI workflows. The hype tells you everything needs a massive, general-purpose generative model. The reality of production tells you that high-frequency operational loops require fast, cheap, and reliable multiple-choice questions.
Of course, the open-source community didn’t just sit back. Harsha Gundal replicated the logic using Qwen 2.5 1B, stripping out text generation and just reading candidate probabilities. People are running these micro-judgment models locally on MacBooks. The magic isn’t in a proprietary black box; it’s in the architectural shift.
But let’s be clear on the boundary. When a task requires the model to generate new parameters, write a line of code, or create text that doesn’t exist in a predefined list, Jev’s advantage vanishes. You still need a heavy LLM for generative tasks.
The era of ‘one giant model to rule them all’ is ending in production. We are moving to a hybrid stack: heavy models for complex understanding and creation, and specialized micro-models for routing, judging, and filtering. Stop trying to make your LLM do everything. Let it think, and let the micro-models do the work.
FAQ
Q: Isn't this just zero-shot classification? BERT could do this years ago.
A: Conceptually, yes. But the difference is in the native architecture and API design. Jev is optimized to batch multiple-choice questions, return structured probabilities, and operate at a latency and cost point that fits seamlessly into real-time agent loops without the heavy setup of traditional encoders.
Q: How do I know when to use a routing model vs a generative LLM?
A: If the answer exists within a predefined list of options or requires scoring candidates, use a routing model. If the task requires synthesizing new text, writing code, or generating parameters that don't exist in a prompt, use a generative LLM.
Q: Are generative LLMs just a fad then?
A: No, they are the heavy lifters. Generative models are still required for complex understanding and content creation. The contrarian take is that generative LLMs are being misused as universal routers when they should be reserved for actual generation tasks.