The Bus Factor Test: A Quick Heuristic for Finding Your Team’s Hidden Single Points of Failure josedacruz, August 11, 2026August 11, 2026 TL;DR: The “bus factor” is the number of people who could disappear from your team before a system stops working — and for most teams it’s uncomfortably close to one. Once you find a low bus factor, the real decision isn’t whether to fix it, it’s how: documentation works well for stable, simple systems, while cross-training works better for complex, fast-changing ones. Pick based on what the knowledge actually looks like, not on habit. What actually matters here The “bus factor” of a system is a dark little thought experiment: if the person who understands it got hit by a bus tomorrow (or, more realistically, took another job), how much trouble would you be in? If the answer is “a lot,” that system has a bus factor of one. It has a single point of failure, and that point of failure is a person. Most teams don’t find this out by asking. They find out during an incident, when the one person who knows how the payment retry logic works is on vacation, and everyone else is reading old commit messages trying to guess what a function called doTheThing() actually does. The fix people usually reach for is “write more docs.” Sometimes that’s right. Sometimes it’s a waste of a week, because the knowledge that actually matters isn’t the kind that survives being written down. So before deciding how to fix a low bus factor, you need to ask what kind of knowledge you’re actually protecting. Three things matter most: How complex is it? A well-documented API endpoint is simple to hand off. A tangled reconciliation script full of special cases learned the hard way is not — no document fully captures “and then I noticed the timezone bug in March, so I added this check.” How much does it matter if it breaks? A low-traffic internal tool failing for a week is annoying. A payment system failing for an hour is a headline. How fast does it change? Documentation for a system that changes weekly is out of date within a month. Documentation for a system that’s been stable for two years is still accurate two years later. Plotting a system on complexity and impact gives you a rough map of where your risk actually concentrates, and a first hint at what kind of fix fits: Where a system lands on this map is a better guide than “just write it down” for everything. Two ways to raise your bus factor Once you know a system is risky, you’re really choosing between two different tools, and they’re not interchangeable. Documentation means writing down what the system does, why it does it, and how to operate it — a runbook, an architecture doc, comments in the code. It’s fast to produce and cheap in the moment. Its weakness is that it only captures what the author remembers to write down, and it goes stale the moment the system changes and nobody updates the doc. Cross-training means having a second person actually work on the system — pairing, rotating on-call, letting someone else make the next three changes instead of the expert doing it alone. It’s slower and it costs real time upfront, because two people are now doing what one person used to do faster alone. But the knowledge that ends up in a second person’s head includes the stuff that never makes it into a doc: the judgment calls, the “don’t touch this on a Friday” instincts, the shape of the system that only shows up once you’ve debugged it yourself. Documentation is the cheaper first move. Cross-training is the one that actually survives contact with a complex system. Neither approach is “correct” in general. The question is which one matches the kind of knowledge you’re trying to protect. A concrete example Say you’re on a small team that runs an online store. There’s a script that reconciles payments between your checkout system and your payment processor every night, quietly fixing small mismatches before they turn into customer complaints. One engineer, Maria, wrote it eighteen months ago and has touched it every time something weird happened since. Nobody else has ever opened the file. Someone runs the bus factor test: “If Maria left tomorrow, could anyone else run this?” The honest answer is no. So now there’s a decision to make, and it’s tempting to default to “let’s get her to write a doc this week.” But look at what this system actually is. It’s complex — full of edge cases learned from real payment failures, not from a spec. It’s high impact — if it breaks silently, money goes missing and nobody notices until a customer complains. And it changes moderately often, every time the payment processor tweaks its API. That combination — complex, high-impact, evolving — is exactly the profile where documentation quietly fails. A doc written this week will describe the system as it is this week. Six months from now, after two more edge cases get patched in, the doc will be dangerously out of date, and whoever’s holding it will trust it anyway. So the team picks the other tool. For the next month, a second engineer pairs with Maria on every change to the reconciliation script instead of her just fixing it solo. It’s slower in the short term. But three months later, when Maria is out during an actual incident, the second engineer has the judgment to handle it — not because they read a doc, but because they’ve already debugged this exact kind of problem once before, with Maria next to them. A lightweight runbook still gets written afterward, covering the basic “how to run this and where the logs are.” That part is simple and stable enough that documentation works fine for it. The two tools aren’t competitors — they cover different parts of the same problem. A rule of thumb You won’t get to fix every single point of failure on a team — there isn’t enough time, and not every risk is worth fixing. Bus-factor risk tends to cluster in a handful of predictable places, and it’s worth knowing roughly where before you go looking for it: Illustrative breakdown — the exact split varies by team, but risk rarely spreads evenly. When you do find a low bus factor worth fixing, use this as your default: reach for documentation first when the system is simple and stable — it’s cheap, and it’ll stay accurate for a while. Reach for cross-training when the system is complex, high-impact, or changes often — because the knowledge that matters there is exactly the kind a document can’t hold. And when something is both simple and low-impact, it’s fine to just accept the risk and move on. Not every single point of failure is worth the cost of fixing. The bus factor test itself takes thirty seconds to ask about any system: “if this person disappeared tomorrow, who else could do this?” The hard part was never asking the question. It’s choosing the right fix once you don’t like the answer. Related architecture bus factorDecision Makingdocumentationknowledge sharingteam resilience