The Hidden Tax That’s Killing the Open Web

Imagine paying for someone else’s dinner. Every single day. That’s exactly what’s happening to millions of small website owners right now.

My friend runs a modest blog about vintage cameras. Last month, his hosting bill tripled. Not because of traffic. Not because of a viral post. Because of bots. Specifically, the bots that AI companies send to scrape every word, every image, every line of code from his site—for free.

Last week, a court ruled that Google and Reddit don’t own the internet. A web scraper won. But the real story isn’t about legal ownership. It’s about who’s left holding the bill.

Every time an AI company scrapes your site, you’re not just losing content—you’re paying for their compute.

Let’s be clear: this isn’t about copyright. The provocative angle that nobody’s talking about is infrastructure taxation. Small website owners are effectively subsidizing the data and compute costs of trillion-dollar AI corporations. Your bandwidth, your server resources, your hosting fees—all consumed by machines that never click a link or buy a product.

I’ve seen this firsthand. A site with 10,000 monthly human visitors can suddenly get 50,000 bot requests in a day. The hosting company doesn’t care if it’s a person or a script. You pay per gigabyte, and the robots are insatiable.

Meanwhile, the AI companies argue that scraping is necessary for innovation. That the open web belongs to everyone. But if that’s true, why are they the only ones allowed to take without giving back?

The open web isn’t dying. It’s being bled dry by trillion-dollar companies who treat your server like a self-service buffet.

You’ve probably noticed the change. Websites are putting up paywalls. Blocking unknown user agents. Locking content behind logins. The walled gardens are growing not because creators want them, but because they can’t afford to feed the bots.

This is the twist: the very thing that made the internet great—open, decentralized, free—is being weaponized against its creators. AI companies rely on that openness to vacuum up data, then sell the results behind closed APIs. They take the commons and privatize the output.

And small creators? They’re left with the choice: close your doors, or pay the hidden tax.

I’m not saying all scraping is bad. But when a single bot from a multi-billion dollar AI lab can cost a hobbyist site $200 a month in extra bandwidth, something is broken. Neutrality is death here. I’m taking a side: the current system is a shakedown, not a partnership.

So what do we do? Some propose blocking all bots. Others want a scraping tax—a per-request fee enforced by the hosting providers. Maybe the solution is turning the tables: charge AI companies for every byte they consume, the way you’d pay for milk at a grocery store.

The real question isn’t whether scraping is legal. It’s whether the open web can survive being treated as a free resource. The next time you find a valuable piece of information online, ask yourself: who paid for the server that delivered it? Chances are, it wasn’t the AI company that just used it to train their model.

This isn’t just a tech issue. It’s a fairness issue. And if we don’t solve it, the web you love will become a gated collection of paywalled silos—where only the big players can afford to play.

FAQ

Q: Isn't web scraping just fair use?

A: Fair use is a legal defense, not a free pass. The issue isn't copyright—it's infrastructure cost. AI companies are using your server resources without compensation. That's not fair use; it's freeloading.

Q: What does this mean for my website?

A: If you're not blocking bots, you're likely paying for AI training data. Check your hosting analytics for requests from known AI crawlers (GPTBot, CCBot, etc.). Consider blocking them or charging per request via a service like Cloudflare's Bot Management.

Q: Shouldn't small sites welcome scraping for exposure?

A: That's a myth. Bots don't click links, don't buy products, and don't generate ad revenue. They drain your bandwidth and return nothing. Exposure is a lie when the exposure is invisible to humans.

📎 Source: View Source