Skip to content
José da Cruz

Enterprise Architect, and Life Thinker

José da Cruz

Enterprise Architect, and Life Thinker

Normalization of Deviance: How ‘It Worked Last Time’ Becomes Your Next Outage

josedacruz, August 11, 2026August 11, 2026

TL;DR: Normalization of deviance is what happens when a team skips a safety step, nothing bad happens, and the skip quietly becomes “how we do things” until it eventually causes a real outage. This is a practical playbook for catching that drift early — name shortcuts out loud, track near misses instead of only failures, and treat “it worked last time” as a warning sign, not proof you’re safe.

Here’s a pattern that shows up in almost every serious outage postmortem: nobody broke a rule on the day things went wrong. The rule had already been broken, quietly, many times before — and each time nothing bad happened, so the broken version became the rule.

This idea has a name: normalization of deviance. A sociologist named Diane Vaughan coined it while studying the Space Shuttle Challenger disaster, where engineers had flagged a known O-ring problem for years. Each launch that didn’t explode made the problem feel a little more acceptable, until one day the odds caught up. You don’t need a spacecraft to see this pattern. It happens constantly in software teams, just with smaller stakes and less dramatic footage.

Picture a mid-size product team with a release checklist that includes a full load test before anything ships to production. One sprint, the team is up against a deadline, and someone says “let’s skip the load test just this once, the change is small.” Nothing breaks. Next sprint, under similar pressure, someone remembers that it was fine last time, so they skip it again. A few months later, skipping the load test isn’t an exception anymore — it’s just how releases work now. Then one day a “small” change turns out to interact badly with a traffic spike, and the site falls over during a promotion. The checklist still existed. It just hadn’t meant anything for months.

The dangerous part isn’t the first skip. It’s that skipping worked, and “it worked” quietly got treated as evidence that the step was unnecessary all along. Here’s a playbook for keeping that drift from happening on your team.

1. Name the shortcut out loud

The first skip is rarely a deliberate policy change. It’s a quiet, individual decision made under pressure, and it often isn’t even discussed as a decision. The fix is cheap: whenever a safety step gets skipped, say so, in words, somewhere the team can see it. “Skipping the load test for this release because of the deadline” in a Slack channel or a PR description turns a silent slide into a visible, reviewable choice. It also means the second and third time it happens, someone can actually notice a pattern instead of each skip feeling like its own isolated, forgettable event.

2. Track near misses, not just failures

Most teams only formally review what they do after something breaks. But the load-test-skipping team had several releases where skipping the test was, honestly, lucky — the change happened to be safe, but nobody actually confirmed that in advance. A near miss is when a shortcut didn’t cause damage this time, but it easily could have. If your team only writes postmortems for outages, you’re only learning from the cases where the dice landed badly. Keep a lightweight log of near misses too, even a single shared doc. Over time it becomes very obvious which shortcuts are getting away with it purely on luck.

Timeline showing a safety step eroding across releases until an outage
The drift is rarely one big decision — it’s a series of small, individually reasonable ones.

3. Make the safety step louder than the deadline

Safety steps lose to deadlines when the cost of skipping them is invisible and the cost of keeping them is obvious. If the load test takes two hours and the deadline is in one, the math always favors skipping — unless skipping has its own visible cost. Some teams solve this by making the check part of the deploy pipeline itself, so skipping it requires an explicit override rather than just quietly not running it. Others require a second person to approve any skip, which is a small amount of friction that makes the shortcut feel like the exception it actually is, instead of a private judgment call.

4. Rotate who gets to say “just this once”

If the same one or two people keep approving shortcuts, they build up a personal track record of “it worked when I allowed it,” which makes the next approval easier and easier to justify — to themselves. Rotating who has the authority to approve an exception means fresh eyes keep re-asking the basic question: is this actually safe, or are we just used to saying yes? It also spreads the accountability, so no single person quietly becomes the team’s designated rule-bender.

Quadrant chart classifying shortcuts by how often they happen and how bad the consequence would be
Not every shortcut deserves the same scrutiny — the ones in the top right are the ones that eventually bite.

5. Re-litigate old exceptions on a schedule

An exception that was reasonable six months ago might not be reasonable now. Maybe the load-test skip made sense when the product had a hundred users and a predictable traffic pattern. It stopped making sense the day the product got featured somewhere and traffic became spiky and unpredictable — but nobody ever went back to check, because the exception had already blended into “how we work.” Put a recurring calendar reminder, even quarterly, to look at your list of standing shortcuts and ask whether the original reasoning still holds.

6. Treat “nothing broke” as inconclusive, not proof

This is the mental model underneath all the other rules. “We skipped it and nothing bad happened” feels like evidence the step was unnecessary. It usually isn’t. Most safety checks exist to catch rare, expensive problems — by design, they’re going to look useless most of the time, right up until the day they aren’t. The absence of a bad outcome after one skip tells you almost nothing about the odds of a bad outcome after the tenth skip. Treat “nothing broke” as “we got away with it,” not “we didn’t need it.”

Pie chart of outage postmortem root causes, with skipped safety checks as the largest share
In a lot of real postmortems, “we knew about this and had stopped checking” is the single biggest slice.

The takeaway

Normalization of deviance isn’t caused by one reckless decision. It’s caused by a long series of small, individually reasonable ones, each one leaning on the last one’s luck. You don’t stop it with a stricter rule — you stop it by keeping the drift visible: naming shortcuts when you take them, logging near misses, spreading out who can approve exceptions, and periodically asking whether last year’s reasonable compromise is still reasonable today. The checklist you stopped following is still a checklist. It just isn’t protecting you anymore.

Related

architecture engineering cultureincident responsepostmortemsrisk managementsystems thinking

Post navigation

Previous post
Next post

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

  • architecture (154)
  • artificial-intelligence (6)
  • books (2)
  • courses (4)
  • decision-frameworks (4)
  • finances (1)
  • java (8)
  • life-improvement (20)
  • metrics (3)
  • observability (2)
  • puzzles and challenges (3)
  • reviews (1)
  • security (8)
  • spring-boot (36)
  • Systems Thinking (3)
  • GitHub Weekly: MCP Tools Trending Right Now (September 28 – October 4, 2026)
  • The Bulkhead Pattern: Why One Slow Dependency Shouldn’t Sink the Whole Ship
  • Essential Reading: A Philosophy of Software Design — Fighting Complexity One Module at a Time
  • GitHub Weekly: Top 10 Trending Repos (September 21 – September 27, 2026)
  • The Volunteer Tax: Why Raising Your Hand in a Meeting Makes You the Permanent Owner
©2026 José da Cruz | WordPress Theme by SuperbThemes