Resources · Patterns

Parallel run and shadow traffic: proving a rebuild before customers touch it

In short

Green tests and a good demo don't prove a rebuilt workflow behaves like the system it replaces. A parallel run does: feed both the same real inputs by replay or shadow traffic, keep side effects such as emails, payments and webhooks contained, compare outcomes, data, permissions, performance and report totals, and make the go/no-go call from a proof report rather than from assurances.

Abstract illustration: two parallel lanes of requests compared side by side

The rebuilt workflow is finished. Tests are green, the demo went well and the team is confident. Now someone has to decide whether real customers, real invoices and real permissions move onto it. If you've lived through one rewrite that looked fine until the first month-end close, you know how little a demo proves.

There is a better basis for that decision. Run the current system and the rebuild against the same real inputs, side by side, and measure where they disagree. Engineers call this a parallel run, and there is a decade of prior art on doing it safely. Five decisions come up:

  1. Replay or shadow — how the rebuild is fed the same inputs.
  2. Side effects — how to stop it emailing, charging or calling anyone while under test.
  3. What to compare — outcomes, data, permissions, performance and report totals.
  4. What "match" means — separating real differences from noise.
  5. The go/no-go call — reading a proof report and acting on what it shows.

Let's see why each of them matters.

Why green tests are not proof

Tests written for the rebuild check what the team believes the system should do. A legacy system does what it actually does. Michael Feathers' characterization tests, which record observed rather than intended behavior, are the right idea, but nobody can write by hand every rule, exception and odd record that fifteen years of production have accumulated.

GitHub met this when it replaced its permissions code. Its engineers wrote that "given enough time and volume, your data will have bugs, too," and that behavior with production data as the input "is the only true test of its correctness compared to the legacy system's behavior." Their answer was Scientist, a library that runs the old code (the "control") and the new code (the "candidate") on every call, returns the old result, and records every mismatch and timing difference. Production traffic became the test suite.

1. Replay or shadow

Both start from the same place: the current system keeps serving and remains the system of record. The question is how the rebuild sees the same inputs.

Replay

Record requests, jobs and data changes from production and play them back through the rebuild against a copy of the data as it stood. Replay is repeatable: the same month of invoicing can go through the twin ten times, and it is the only option for infrequent work such as a quarterly rebate run. The cost is fidelity: state has to be restored to the right point, and anything fetched from a third party at the time has to be stored or stubbed.

Shadow traffic

Copy live requests to the rebuild as they arrive and discard its responses. Martin Fowler describes this under dark launching: old and new code are both called and their results checked, "but only one answer returned to the interface." Istio's traffic mirroring, for example, sends "a copy of live traffic to a mirrored service" outside the critical request path. Shadowing gives you real load, real timing and the ugly records nobody put in a test fixture, at the price of repeatability and much stricter control of side effects.

In practice you want both: shadow the high-volume, read-heavy paths for breadth and performance, replay the stateful, periodic and rare paths for depth. GitHub uses Scientist "only on read operations," and side effects are the reason.

2. Controlling side effects

A rebuild under test must never act on the world, and there are more ways it might than most teams expect:

The mechanism matters less than the discipline. Every outbound path goes through an adapter set to live, sandbox or record-only, and during proof none is live. Your engineers can verify that before the first request is copied.

3. What to compare

Comparing HTTP responses is the least interesting layer. The proof should cover every output a workflow produces:

4. What "match" should mean

Compare raw output byte for byte and you'll learn nothing. Timestamps, generated IDs, list order and floating-point rounding all differ between two systems that are behaving identically. Twitter built Diffy around that fact. It sends each request to the candidate and to two copies of the known-good code, and whatever differs between the two good copies is treated as noise. Whatever the tool, run the current system against itself first to learn what "same" looks like.

From there, "match" becomes a written rule per field, agreed before the comparison starts: monetary totals match exactly, timestamps within a tolerance, free text ignored unless it carries business meaning. Those rules belong in the specification, so a disagreement about what counts is a specification question, not an argument.

Every remaining difference is then classified. In our experience they fall into four groups: a bug in the rebuild (the common case), a bug in the current system that customers rely on, a data-quality problem neither system handles well, and a gap in the specification nobody wrote down. Only the first is simply "fix it". The other three need an owner to decide, and the decision belongs in the record.

5. Reading the proof report and making the call

A proof report should let a non-engineer answer three questions. How much of the specified behavior has been exercised? Where do the two systems still differ? Who owns each difference, and what was decided? If it can't answer those, it's a log, not a report. Coverage matters as much as match rate: a 100% match on 30% of the specified behavior is not a pass.

A publisher's subscription system we worked on makes the point. The rebuilt renewal workflow matched on well over 99% of a month's replayed renewals. The remainder was all accounts created before a pricing change years earlier, where the old code prorated refunds with a rounding rule that lived nowhere but in the code. Small in count, concentrated in the longest-standing customers, and invisible in any demo. The owner chose to keep the old rule, write it into the specification and run the comparison again.

The go/no-go call then becomes concrete:

"We're at 99.7%, can't we just go?" is the objection you'll hear. It depends on what's in the 0.3%. A formatting difference in an internal log, yes. Twenty invoices for your largest accounts, no, and the report should make that obvious.

What a parallel run buys you

A parallel run makes the question "is it correct?" answerable with figures instead of assurances, while the current system is still carrying the business. The steering meeting moves from "do we trust the team?" to "what is in the 0.3%, and who decided?"

Questions to ask whoever proposes to move traffic:

Where Fabrica stands

Attest, the fourth stage of Fabrica's CLEAR method, applies this practice to a locked System Specification. Representative workloads are replayed or shadowed with side effects controlled, and the rebuilt workflow and the current system are compared on business outcomes, data, reports, permissions and performance. Differences are tracked as fixes with an owner, and the output is a proof report to support a go or no-go decision, which your engineers can check before any cohort moves.

Keep reading

Related

Which priority is your system holding up?

Tell us about it on a 30-minute call. We'll suggest a first scope and what a technical review would need to confirm.

Book a 30-minute call