Resources · Decision guide

Key-person risk: when two people hold the whole system in their heads

In short

Key-person risk rarely reaches a risk register because it is a condition, not an event, and no single function owns it. Measure it per workflow with the bus factor, recover the knowledge from the running system and owner review rather than interviews alone, and run retention, documentation and succession as one plan. The first 90 days decide most of it.

Abstract illustration: two people carrying a system others cannot reach

Ask who could change the pricing rules without breaking month-end, and the answer is a name. Sometimes two. They are reliable and long-serving, and nobody worries until one of them books a meeting with HR.

If that describes your core system, you carry a risk that is probably not on your risk register. It isn't a people problem that HR owns or a technology problem that IT owns. It sits between the two, which is why it so rarely gets written down.

The takeaway: treat key-person risk as a property of the system, not of the people. Measure it per workflow, recover the knowledge from the running system rather than from interviews alone, and run retention, documentation and succession as one plan. The first 90 days decide most of it.

Why it rarely shows on a risk register

A risk register lists events with a likelihood, an impact and an owner. Key-person risk is a condition, not an event. "Marta resigns" reads like gossip, not a line item, so it never gets a row.

It also has no natural owner. HR sees a retention question, IT sees a documentation gap, and finance sees nothing until a month-end close doesn't reconcile. Each function handles its slice, and the risk as a whole belongs to no one.

The concern is widespread even so. In a 2022 survey of 1,000 IT managers, 67% were concerned about losing knowledge when people leave, 64% said it had already happened to them, and only a quarter had a knowledge management strategy. The worry exists. The register entry doesn't.

How to measure it: the bus factor

The software world has a blunt name for this: the bus factor, also called the truck factor. It is the smallest number of people who would have to leave before work on a system stalls. A bus factor of one means a single resignation stops you changing the system safely.

The research is sobering. Avelino and colleagues estimated the truck factor of 133 GitHub projects and found that 65% had a truck factor of two or lower. These are public projects with large contributor communities; an internal application maintained by six people is unlikely to score better. A follow-up study of 1,932 projects found that 16% were abandoned after their core developers left, and only 41% of those survived.

Two more findings shape what to do about it. A 2022 survey of 269 engineers at JetBrains found that 63% had worked on at least one project where they felt there was a high risk of the bus factor reaching zero. And Rigby and colleagues, studying Chrome and Avaya, found projects exposed to knowledge losses more than three times the expected loss. Plan for the bad quarter, not the average one.

A simple assessment you can run this week

Don't score the whole system. Score the workflows the business cannot run without, one at a time.

Why interviews alone won't recover the knowledge

The usual response is to interview the key people. It is necessary, and it is not enough. Michael Polanyi put the reason in one line in 1966: "we can know more than we can tell." Experts hold much of what they know as a feel for the system and cannot recite it on request.

Two failure modes follow. People describe what the system was designed to do, not what it does after twelve years of patches. And they don't know what they know until a concrete case prompts it.

We saw this with a publisher's subscription system. The owner described three renewal rules, and described them well. The logs for one quarter showed eleven distinct paths, including a grace-period rule for a single enterprise account. Nobody had misled anyone; we had asked the wrong kind of question.

Michael Feathers made the same point about code with characterization tests, which pin down what the code does rather than what it should do. Observe first, then ask.

Recovering knowledge from the running system and owner review

In our experience, three sources work, in this order. The running system shows what happens. Code, schema and reports show how. Owners confirm why, and whether it is intended.

Observe real behavior first

Record what the current system does in production: requests, background jobs, queries, errors and report runs, tagged by release. Make sure the window covers the rare paths, such as month-end or the regulator's report. Anything not yet observed is recorded as a gap, not assumed. With that evidence beside you, reading the code becomes a search for the branch that produced a specific outcome, not a tour of 400,000 lines.

Review with owners, using evidence

Now bring in the key people, but change the question. "Last Tuesday, 41 renewals were processed at a discount that isn't on the price list. Is that intended?" gets a precise answer in two minutes. Their scarce hours go into confirming and correcting, where their knowledge is worth most.

Write rules as testable requirements and track coverage

Each confirmed rule becomes a plain-language requirement with a stable ID, a link to its evidence and a named owner. Keep a running view of what is covered and what is still a gap; coverage is the measure that matters, not page count.

Retention, documentation and succession as one plan

Most organizations run these three separately, and they undermine each other. A retention bonus on its own buys a few months and signals that the knowledge is the person's asset to hold. Documentation on its own, assigned to the key person as a side task, never finishes. "We don't have time to document — the deadline is next quarter." Succession on its own, hiring someone to shadow the expert, fails because there is nothing to shadow but habits. Run them as one plan with one owner, usually your engineering lead with an executive sponsor.

What to do in the first 90 days

  1. Days 1 to 30: map the exposure. Run the assessment above and pick the three workflows with the highest cost if they stall. Talk to the key people plainly; they usually know they are a single point of failure and often find it a burden. Agree the retention window and set up read-only observation.
  2. Days 31 to 60: capture and review. Let evidence accumulate, including at least one month-end. Hold owner reviews on specific cases, two hours a week rather than two days a quarter. Write each confirmed rule as a requirement, record the gaps, and name a second owner per workflow.
  3. Days 61 to 90: prove the transfer. The second owner makes a real production change to each workflow, with the key person reviewing rather than doing. Then decide whether to keep the system, upgrade it in place or renew it by workflow, now with a specification to decide from.

Ninety days will not finish the job for a large system, but it moves your three most exposed workflows off a bus factor of one and shows you what the full plan costs.

Where Fabrica stands

Fabrica's CLEAR method starts from this problem. Capture instruments the running system and records how it behaves, with unobserved paths marked as gaps, and Lock turns that evidence and the owner reviews into a System Specification with stable IDs and owner sign-off that stays current as linked work completes. Our hypothesis is that observation plus owner review recovers more of the rules, in fewer of the key people's hours, than interviews alone; a pilot on one workflow measures that. See how it works and the IT and risk page.

Start before the resignation

Key-person risk is cheapest to address while the key people are still in the building and not yet looking around. The work is not glamorous: observation, specific questions and a second name on each workflow. But it turns a risk that lives in two heads into one that lives on paper, with a number, an owner and a plan.

Questions to ask about your own system:

Keep reading

Related

Which priority is your system holding up?

Tell us about it on a 30-minute call. We'll suggest a first scope and what a technical review would need to confirm.

Book a 30-minute call