You know the exact feeling. You’ve written a clean, perfectly logical function. It does one thing, and it does it well. You push it for review, and within seconds, the CI pipeline spits back a red X. Your “Change Risk Anti-Pattern” score is too high. The machine has deemed your code CRAP.
So, what do you do? You don’t rewrite the logic. You don’t improve the design. You slice that function into three awkward, fragmented pieces that don’t make sense in isolation. You pass the metric, merge the code, and silently apologize to the next developer who has to read it.
We didn’t eliminate the crap; we just taught it to wear a tuxedo.
Back in 2011, the Google Testing Blog introduced the concept of CRAP—Change Risk Anti-Patterns. The idea was simple: use cyclomatic complexity and test coverage to calculate a risk score. High score meant high risk. It was a well-intentioned heuristic meant to flag spaghetti code before it metastasized. But like every good intention in software engineering, it got paved over by bureaucratic enforcement.
We took a risk heuristic and mistook it for an objective verdict. We turned a suggestion into a hard quality gate.
When human judgment surrenders to a threshold, the threshold becomes the product.
The irony is thick enough to choke on. The metric named “Change Risk Anti-Pattern” has itself become a change-risk anti-pattern. By trying to codify quality, we invited superficial refactors that make complexity scores look pristine while the actual design debt rots in the background.
You see it every day. Classes split into meaningless abstractions just to keep the per-method line count down. Cyclomatic complexity dodged by burying logic in strings or polymorphic explosions that are ten times harder to debug. The scanner gives you a green checkmark, but the architecture is bleeding out.
A metric is just a map; it will never tell you where the landmines are buried.
The darkest secret of modern code review isn’t that developers write bad code. It’s that developers are actively gaming the metrics to appease bots. The tool created to identify crap has become the most reliable generator of crap, because it optimizes for the map, not the terrain. It doesn’t just catch bad code—it creates a new, mutant species of it, engineered solely to fool the scanner.
If you’ve ever felt your soul leave your body while refactoring a perfectly fine method just to satisfy a SonarQube warning, you aren’t crazy. You are a victim of a system that values easily measurable nonsense over actual engineering judgment.
It’s time to call this what it is. Complexity metrics are not a substitute for taste. They are a crutch for managers who don’t trust their engineers. If your quality gate forces you to write worse code to pass the check, your quality gate is broken.
Stop worshipping the numbers. Start trusting the humans who actually have to maintain the terrain.
FAQ
Q: Aren't complexity metrics better than nothing for enforcing baseline standards?
A: No, they're often worse than nothing because they create the illusion of quality. 'Nothing' forces you to actually look at the code. A bad metric lets you rubber-stamp garbage just because the scanner gave it a green checkmark.
Q: So should we just delete all our CI quality gates?
A: Delete the hard blocks. Keep the warnings as informational data points, not merge blockers. If a metric flags high complexity, let a human review it and decide if it actually matters in context.
Q: If developers just wrote simple code, we wouldn't need these metrics anyway, right?
A: If managers just trusted developers to write simple code, we wouldn't have developers splitting functions into fragmented abstractions just to appease a bot. The metric creates the exact behavior it claims to prevent.