Resources · Engineering

Your production system is the specification: rules from observed behavior

In short

Reading the code and interviewing owners each recover about half of a legacy system's rules, and not the same half. The running system holds both. Derive requirements from observed behavior using instrumentation, analytics, schema and report lineage, collected read-only and redacted. Write each rule as a testable requirement with a stable ID, and track coverage and known gaps before anyone approves a specification.

Abstract illustration: a live traffic waveform turning into a written specification

You've decided to do something about the system that runs your business. Before anyone writes new code, someone asks the obvious question: what exactly does the current system do? The documentation is thin or years out of date, and the people who built it have mostly moved on.

Most teams answer with one of two methods. Engineers read the code and write up what they find, or analysts interview the people who use the system and write up what they say. In our experience each method recovers about half of the rules, and not the same half.

The takeaway of this piece is simple. The specification you're looking for already exists. It's your production system, running right now, and every rule it applies is observable. Treat the running system as the specification, derive requirements from what it does, and write them in a form owners can approve and engineers can test.

Why code reading and interviews each miss half

Michael Feathers defined legacy code as code without tests. The definition names the real problem. Without tests, nobody can say what the system is supposed to do, only what it currently does, and the two drift apart over the years.

What the code can't tell you

Code shows what could happen, not what does happen, how often, or for whom. A branch that handles a discount rule may be the heart of your pricing or dead since 2017. The code looks the same either way.

Reading code is also slow. A field study of 78 professional developers found they spent about 58 percent of their time on program comprehension. A large codebase means months of reading before anyone writes a rule down.

Then there is Hyrum's law: "With a sufficient number of users of an API, it does not matter what you promise in the contract: all observable behaviors of your system will be depended on by somebody." A sort order nobody specified. A date format a partner's import script parses. None of these are in the code as rules. They're there as accidents, and your customers depend on them.

What the owners can't tell you

Interviews recover intent: why a rule exists and what the business wants from it. What they can't tell you is what the system does in cases they've never seen. Ask a finance lead how invoicing works and you'll get the happy path. Ask about an allocation with no rate card and you'll get a pause.

Owners also describe the process as it should be, not as it is. The exception a support agent fixes by hand every month-end is part of the system's behavior, and it rarely comes up in a workshop. The person who knew the oldest rules left years ago.

Code gives you the mechanism without the usage. People give you the intent without the edge cases. The running system holds both.

Characterization tests, at the scale of a system

Feathers's answer for untested code was the characterization test: run the code, record what it does, and turn the recording into a test. The test doesn't claim the behavior is correct. It pins the behavior so you'll know when it changes.

Deriving requirements from production is the same idea one level up. Instead of calling a function and recording its return value, you observe the live system: the requests it serves, the jobs it runs, the rows it writes, the reports it produces. Each observed behavior becomes a candidate rule, and owners decide whether it's intended, accidental or wrong. The judgment stays with people. The inventory comes from the system.

Where the evidence comes from

No single source is enough. Five sources cover most of a web application's behavior and correct each other's blind spots.

Owner review sits across all five, and evidence changes the conversation. "On the 14th, 312 invoice lines were created with a zero amount. Is that intended?" gets a better answer than "how does invoicing work?"

Observing production without putting it at risk

"Our security team won't let anyone near production." The concern is fair. The answer is to agree the rules of observation before the first byte is collected.

Write what you see as testable requirements

Observed behavior isn't a specification until it's written in a form that can be approved and tested. "The system handles unbillable allocations correctly" can't be tested, because nobody can say what would make it fail.

We prefer a structure close to EARS, the Easy Approach to Requirements Syntax that Alistair Mavin and colleagues developed at Rolls-Royce and published in 2009. Each requirement names a condition, a system and a response, in that order:

REQ-c3d4: While an allocation has no rate card, the system shall flag the invoice line as unbillable rather than invoicing at zero.

"While" names the state, "the system shall" names the response, and a tester can set up the condition and check the result. Three more things make a rule useful over time:

Each requirement also carries a status: observed, confirmed by an owner, rejected as a defect, or intended but not yet observed. That last one matters more than it looks.

Track coverage and name the gaps

How do you know when you've seen enough? You don't, not completely, and pretending otherwise is how projects discover a rule in week 40. The practical answer is to track coverage against a threshold agreed for the scope, and to record whatever hasn't been observed as a gap, in writing, with an owner.

How long observation takes depends on three variables: traffic volume, the length of your business cycles, and the share of rare paths. A high-volume checkout flow can show most of its behavior in days. Anything tied to month-end or annual renewal needs at least one full cycle. On a publisher's subscription system we worked on, a quarter of observation showed almost no renewal traffic for annual plans, because most of them renewed in a single month. The gap was recorded, the first scope was narrowed to monthly plans, and the annual cohort waited for a renewal cycle.

A gap is not a failure. It's a decision point with three options: capture for longer, narrow the scope, or accept the risk explicitly with a rollback plan. What you must not do is let a gap pass as coverage.

Where Fabrica stands

Fabrica's CLEAR method applies this directly. Capture connects instrumentation, product analytics, schema and report lineage into an evidence map, with payloads redacted and collection off the request path. Lock turns observed behavior and owner review into requirements with stable IDs, which make up a versioned System Specification approved when coverage and confidence reach the agreed level. Anything not yet observed is recorded as a gap. More on how it works and engineering.

Start from the system you have

In the 2024 Stack Overflow Developer Survey, 62 percent of developers named technical debt as their top frustration at work. Part of that debt is simply that nobody can say what the system does. The cure isn't a longer workshop or a bigger reading assignment. It's evidence from production, written as rules owners can approve and engineers can test, with the gaps named. Start from what the running system does, and make the document answer to it.

Questions to ask before you approve a specification:

Keep reading

Related

Which priority is your system holding up?

Tell us about it on a 30-minute call. We'll suggest a first scope and what a technical review would need to confirm.

Book a 30-minute call