Your Compression Benchmarks Are Lying. bzip3 Just Proved It.

You’ve seen this movie before. A brilliant developer builds something genuinely impressive. The benchmarks look stunning. The community shows up, curious. And then, within 72 hours, the whole thing collapses — not because the code was bad, but because the credibility was.

Welcome to bzip3.

Let’s be clear about something right up front: the best compression algorithm in the world dies in obscurity if nobody trusts its benchmarks.

Here’s what happened. bzip3 showed up claiming to be “stronger than bzip2.” Four times smaller than zstd in some benchmarks. Impressive, right? The kind of numbers that make you want to rip out your current compression library and swap it in immediately.

But then the community started doing what technical communities do: they looked under the hood.

And what they found was a masterclass in self-sabotage.

The block size for bzip3 was set to 512MB. The window size for zstd? Left at its default — 8MB. On a corpus made of concatenated Perl source code. In other words, bzip3 was given a sledgehammer while zstd was handed a scalpel, and then someone printed a poster declaring the sledgehammer the winner.

Cherry-picked benchmarks don’t prove your tool is better. They prove you’re scared it isn’t.

But it gets worse. The parallel decompression comparison? bzip3 was benchmarked against plain bzip2 — not pbzip2, the parallel version. That’s like racing a Ferrari against a bicycle and then holding a press conference about your victory lap.

And then there’s the build. The latest release is a year old. The last commit was two months ago. The build is failing. Let me say that again: the build is failing. You can’t even compile the thing, and the project is out here claiming dominance over one of the most battle-tested compression tools in computing history.

A failing build is not a bug report. It’s a death certificate.

Here’s the uncomfortable truth that nobody in the compression community wants to say out loud: the bottleneck for a new compression tool was never the algorithm. It was always trust.

Think about it. When you choose a compression library, what are you actually choosing? You’re choosing reliability. You’re choosing the confidence that five years from now, your archives will still decompress. You’re choosing the collective weight of thousands of production deployments, bug reports filed and fixed, edge cases discovered and handled.

bzip2 has been around since 1996. zstd is backed by Meta and battle-tested across petabytes of real-world data. These tools earned their place not through benchmark tables but through a relentless, unglamorous accumulation of trust.

And bzip3 walked into that arena with a failing build and cherry-picked numbers, expecting to be taken seriously.

Trust is the only compression ratio that matters, and bzip3 just compressed its own to zero.

Now, here’s where it gets genuinely painful — because bzip3 might actually be good. The Burrows-Wheeler transform implementation could be legitimately interesting. The technical approach might represent real innovation. We’ll never know, because the author chose to sell it with sleight of hand instead of rigor.

This is the part that should make every open-source maintainer uncomfortable. Your algorithm could be brilliant. Your benchmarks could be honest. Your build could be green. But if you let even one of those slip — if you cherry-pick a single comparison, if you leave a build broken for weeks, if you make a claim like “stronger than bzip2” without defining what “stronger” even means — the community will find it. And they should.

Because here’s what the bzip3 saga really teaches us: in open source, credibility is not a feature you add later. It’s the foundation you build on day one.

The community didn’t reject bzip3’s algorithm. They rejected its methodology. They rejected the implication that developers would be too lazy to check the block sizes. They rejected the assumption that “four times smaller” would be enough to override a failing build and an unfair comparison.

And that rejection — that ruthless, public, uncomfortable rejection — is exactly what makes open source work.

So the next time you see a benchmark that looks too good to be true, check the block sizes. Check the window settings. Check whether the comparison is fair. Check if the build passes.

Because the most compressed thing in any benchmark scandal is always the truth.

FAQ

Q: But what if bzip3's algorithm is actually better?

A: It might be. The Burrows-Wheeler transform implementation could be genuinely interesting. But we'll never know, because the benchmarks were rigged to look impressive rather than to prove anything. When you cherry-pick comparisons and leave your build broken, you make it impossible for anyone to verify your claims independently. A great algorithm buried under dishonest methodology is functionally useless.

Q: How should I evaluate compression benchmarks then?

A: Check three things immediately. First, are the block sizes and window settings identical across all compared tools? If one tool gets 512MB and another gets 8MB defaults, walk away. Second, are you comparing apples to apples — parallel vs parallel, single-threaded vs single-threaded? Third, does the project build cleanly right now? If the build is failing on the main branch, no benchmark number matters.

Q: Is the community being too harsh on a solo developer?

A: No. Open source is a trust market, and ruthless scrutiny is the quality control mechanism. When zstd launched, it survived the same microscope because its benchmarks were fair and its build was green. If you're making public claims about superiority, you're signing up for public verification. Solo developer or corporate team — the standard is the same: show your work, or expect it to be torn apart.

📎 Source: View Source