An Intern Burned 100 Million Tokens in a Day. They Weren’t Fired. They Were Right.

You’ve probably seen the screenshot. An intern burns through 100 million tokens in a single day, outputs exactly 15 lines of code, and gets hauled into a meeting with management. The immediate reaction from anyone with a spreadsheet is panic. The intern must be wildly incompetent, right?

Wrong. The intern might be the only person in the building who actually understands how AI coding works.

You are still using Industrial Era metrics to measure a post-Industrial revolution.

When you hear ‘100 million tokens for 15 lines of code,’ your brain instantly calculates the cost per line. It looks terrible. But this is the fundamental trap of AI Coding ROI. Code generation is cheap. The real cost—the massive, silent budget black hole—isn’t output. It’s input.

Let’s look at what the intern was actually doing. They didn’t tell the AI to ‘write 15 lines.’ They likely told the agent to read a mid-sized project, understand the dependencies, find a deeply buried interface bug, and fix it. To do that, the AI agent has to feed the entire codebase back into the model on every single turn of the conversation. A few files easily hit tens of thousands of lines. Multiply that by ten rounds of trial-and-error, and you’ve hit 100 million tokens before the agent has even written a single semicolon.

AI doesn’t think. It just re-reads everything it forgot, and then pretends to think.

Every time the agent makes a decision, it has to re-input the entire conversation history, the file contents, and the execution results. If you open a new chat window, the context resets to zero. The AI has to read the entire codebase from scratch, re-trace its own steps, and re-make all the same mistakes. That is where your money goes. Not on writing code, but on re-reading, forgetting, and looping.

This is why ‘cost per line of code’ is a completely broken metric. Deleting 500 lines of legacy garbage is infinitely more valuable than generating 50 lines of new boilerplate. But if you measure by lines, deletion is a negative output. A critical one-line bug fix might represent 8 hours of AI exploration, trial, and verification. If that bug fix saves the company a weekend outage, the 100 million tokens (which cost maybe $50 on a fast model, or $400 on a flagship one) is a rounding error compared to the human capital it would have taken to find manually.

So, how do we actually measure this? You stop looking at token price, and you start looking at Cost Per Accepted Task.

Industry benchmarks like SWE-bench prove this. Two agents can have a near-identical pass rate on coding tasks, but one agent’s cost per task can be 70 times cheaper. Why? Because the expensive one loops, forgets, and re-reads. The cheap one manages its context ruthlessly.

Saving tokens locally is the fastest way to burn your budget globally.

GitHub learned this the hard way. They tried to truncate tool outputs to save tokens. The model, starved of context, immediately started re-running commands and re-opening files to figure out what was going on. Total token consumption exploded. Local efficiency caused global waste.

If you want to stop burning money, you need a new accounting system. Here is how you do it:

1. Context Starvation
Stop feeding the AI your entire project directory. If you know the bug is in the payment service, give it the exact file path. Use ignore files to block out node_modules, build logs, and irrelevant directories. Don’t let the AI go fishing. Tell it exactly where to look.

2. Task Slicing
Never give an AI a massive, vague goal like ‘refactor the payment module.’ It will get lost, try something, fail, and loop infinitely. Break it down. Step one: read config. Step two: write migration function. Step three: write tests. Step four: replace references. Small tasks mean small context windows, and small context windows mean cheap iterations.

3. Model Tiering
Not everything needs the $20 flagship model. Code formatting and simple function generation should run on a Flash-level model that costs pennies. Save the heavy artillery for complex architecture and cross-module refactoring. 80% of your daily work can be done by a model that costs 1/150th of the flagship.

4. Memory Injection
If you have to start a new chat, don’t let the AI figure it out again. Paste the confirmed conclusions and key code snippets from the last session directly into the prompt. Better yet, use a memory database to store team knowledge so the AI retrieves architecture maps instead of reverse-engineering them from scratch every time.

At the organizational level, the worst thing management can do is look at a monthly token bill, panic, and implement a blunt budget cap. That is lazy leadership. You are capping the one thing that might be saving you thousands of hours of human labor.

Instead, track three things: consumption (where did the tokens go?), efficiency (how many tasks passed review on the first try?), and value (did we ship faster?).

The intern didn’t waste 100 million tokens. They spent 100 million tokens to solve a problem that would have taken a senior engineer half a day. They traded $400 of compute for $500 of human capital, and they did it while learning the hardest lesson of the AI era:

You can’t cut your way to high ROI. You have to manage what the machine forgets.

FAQ

Q: Isn't burning 100 million tokens for 15 lines of code objectively wasteful?

A: No. If those 15 lines fix a deeply buried interface bug, the tokens weren't spent writing code—they were spent reading, understanding, and verifying. You're paying for the AI's exploration, not its typing speed.

Q: How should we track our team's AI coding efficiency then?

A: Stop tracking 'cost per line of code.' Track 'Cost Per Accepted Task.' Measure how many tokens it takes to get a unit of work that actually passes code review and merges into the main branch.

Q: Should we just put strict token limits on our developers to control costs?

A: That's the worst thing you can do. Blunt budget caps force developers to use weaker models or starve the AI of context, which causes it to loop and fail more often, ultimately burning more tokens globally. Manage the context, not the budget.

📎 Source: View Source