Model Selection

I Spent Months Chasing MTEB Scores. Then I Built Something That Actually Works.

I spent months chasing MTEB scores, only to find my embedding models flopped on my own data. Generic benchmarks are misleading – real performance depends on your specific retrieval pipeline. That’s why I built Embench: a free playground to compare embeddings on your own data and taxonomy. Stop relying on leaderboards. Test on what matters.

You’re Wrong About AI Benchmarks. Here’s What Actually Predicts Your Bill.

Benchmark scores measure how smart an AI model is in isolation, but they completely ignore token consumption β€” the variable that actually determines your bill. Qwen 3.8 and Claude Opus 5 prove that the highest-scoring model is often the most expensive one to run. The real metric that matters isn’t raw performance; it’s cost-per-task for your specific workload.

The New Default That’s Quietly Taking Control of Your Code

Claude Code’s new default auto mode is more than a UX improvementβ€”it’s a quiet transfer of control from developers to Anthropic’s cost-optimization algorithms. The promise of ‘best model for the task’ hides an opaque selection logic that may prioritize cheaper inference over output quality. Developers need to understand the trade-off before they surrender their choice.

The Smartest AI Model Is the One You’re Ignoring

Most LLM benchmarks are designed to sell expensive models, not solve your problems. We tested 8 models on the Baba Is You puzzle game and found that the cheapest model – DeepSeek V4 Flash – often beat the most expensive. Task-specific performance is the only reality. Stop chasing leaderboards, start chasing your actual use case.