You’re building an LLM app, and you’ve run the same prompt through two different models. The token counts don’t match. You shrug—it’s probably fine. It’s not fine. You’re losing money and context window every single time, and nobody told you that the ‘exact prompt’ is a lie.
Here’s the dirty secret of LLM economics: tokenization is not a property of your text. It’s a property of the model’s tokenizer. That means the same string of words can cost you 10% more tokens in one model vs. another—and you’ll never see it coming.
I saw this firsthand when a colleague bragged about optimizing prompts for GPT-4. He’d trimmed every unnecessary character. Then he switched to Claude. The costs shot up. The prompt was ‘exact’ but the tokenizer wasn’t. Every model sees your words through its own lens, and that lens has a price tag.
Take a recent tweet from J.P. Schroeder: ‘Token use for exact prompt across models.’ The top comment? ‘The post does not say what the prompt tried.’ Exactly. Because the prompt itself is irrelevant. What matters is which model’s tokenizer is doing the counting. The moment you assume a universal token count, you’ve already lost.
Think about what that means for your context window. You’re paying for 128k tokens of context, but if your tokenizer inflates each word by 50%, you’re actually getting 85k. You’re not saving money—you’re starving your model of usable space. You’re paying for storage you can’t use.
So what do you do? First, stop believing in ‘exact prompts.’ Second, measure token counts per model, per use case. Build a matrix: your prompt, your tokenizer, your cost. Third, share this with your team because the quietest tax on your LLM budget is the one you never see.
Here’s the twist: optimizing for one model’s tokenizer can silently destroy your performance in another. The harder you optimize for GPT-4, the more you’ll pay when you migrate to Llama. The ‘exact prompt’ is a trap. Tokenization is a fingerprint, not a universal truth. Build for your model, not for the abstraction.
FAQ
Q: Doesn't tokenization work the same way across models?
A: No. Each model uses its own tokenizer (e.g., GPT-4 uses cl100k_base, Claude uses a different one). The same string can produce different numbers of tokens because tokenizers split words, punctuation, and subwords differently.
Q: Can I just use a universal token counter?
A: Only as a rough estimate. For production, you must measure token counts against the exact model you're using. A universal counter gives you a false sense of precision.
Q: Why does this matter if I'm not paying per token?
A: Even if you're not paying, token counts affect context window limits. One model may fit your prompt in its context, another may not. And if you're using an API that charges per token, the difference is direct cost.