Resources · Patterns

Branch by abstraction and feature toggles inside a legacy codebase

In short

Most of a legacy system's business rules sit below the request layer, where a proxy can't route around them. Branch by abstraction puts a seam inside the code, and a cohort toggle turns that seam into a gated 1%, 10%, 50%, 100% release with rollback as a toggle change. The work only ends well if the toggle, the seam and the old code are removed afterwards.

Abstract illustration: a code path forking into toggled routes

You've ruled out the big-bang rewrite. You've read about the strangler fig: put a proxy in front of the current system, build the new version beside it, move one URL at a time. Then your engineering lead points out the problem. The thing holding the business back isn't a URL. It's a pricing engine or an invoice calculator, buried inside the application and called from forty places, including background jobs and a nightly report. No proxy can see it.

This is the normal case. Most business rules in a legacy system live below the request layer, out of a proxy's reach. Incremental replacement still works there, but it needs two things: a seam inside the code, which Paul Hammant and Martin Fowler call branch by abstraction, and a toggle that decides, call by call, which implementation answers. Together they give you the same gated release path a proxy would, including rollback as a toggle change.

The catch is that the seam and the toggle are scaffolding. Leave them up and the "new" system inherits the clutter. So this piece covers four things in order: where the proxy can't reach, how branch by abstraction creates a seam, how toggles and cohorts turn it into a gated release, and how to take the scaffolding down.

When the proxy can't see what you need to replace

A reverse proxy routes on what it can observe: path, host, header, cookie. That works for a workflow with its own URLs and fails for anything shared, which in a ten-year-old Rails, Java or .NET monolith is most of what matters:

You can't route 10% of a nightly invoicing job through a proxy. You can change the job so that 10% of its invoices are calculated by new code. That needs a decision point inside the application.

Branch by abstraction: a seam inside the code

Hammant named the technique in 2007, crediting Stacy Curl, and Fowler's 2014 description is the one most people know. Instead of replacing the old module in one move, you put an abstraction in front of it, move every caller onto it, build the new implementation behind the same abstraction, and switch callers over as the new one proves itself. The system builds and runs at every step. In practice:

  1. Find every caller. Static search misses metaprogramming, and the job nobody remembers surfaces in production.
  2. Introduce an interface for what callers actually need, not what the old module happens to expose. Fowler notes this is the moment to raise test coverage.
  3. Move callers onto the interface one at a time, releasing as you go. Only the plumbing changes.
  4. Build the new implementation behind the interface, on whatever stack you've agreed.
  5. Switch callers over, then delete the old implementation. Hammant's last step, often skipped: remove the abstraction if it has nothing left to do.

The first step is the hard one. On a German fintech client's platform we once counted the call sites of a fee calculation. The team estimated twelve. There were thirty-one, including three in reporting SQL that bypassed the application entirely. A narrower seam would have produced two sets of numbers.

Proving the two implementations agree

An abstraction guarantees both implementations have the same shape, not that they give the same answers. Steve Smith's verify branch by abstraction closes that gap: for a period, the abstraction calls both implementations with the same inputs, returns the old result and records every difference.

For a leader, this is the step that matters, because it's where you learn what the legacy system actually does. Every mismatch is either a bug in the new code or an undocumented rule in the old one. On the fintech platform the comparison ran for several weeks, and the differences were almost all the second kind: rounding at a different step, a currency treated as a special case since 2014, a client override hard-coded in a conditional. None of it was written down. It is now.

Two cautions. Where the new side writes, sends or charges, run it with side effects controlled or you will double-bill someone. And let owners, not engineers, decide what "the same" means: to the cent, or within a reconciled tolerance.

Toggles and cohorts: a release path you can reverse

Once the comparison is clean, the abstraction needs a switch. In Pete Hodgson's taxonomy this is a release toggle: short-lived, evaluated at runtime, and taking some context, usually the customer or account, to decide which implementation answers. That context lets you release by cohort rather than by a random share of requests. Random sampling tells you about error rates. Cohorts tell you about business outcomes, because the same customer sees the new behavior consistently and you can reconcile their invoices end to end. A sensible canary-style path, and the one we'd start from:

Each step is gated on evidence, and the gate is a decision a named person takes, not a date. Rollback at any step is a toggle change: point the cohort back at the old implementation and the running system carries on, with no redeploy. "We'll roll back if something goes wrong" is only credible if rollback was tested before it was needed.

One limit to state plainly. A toggle reverses which code runs, not data the new code has written. If the new implementation owns writes, you need a data transition plan with mappings, backfill and reconciliation agreed before the first live write. That is its own workstream.

Cleaning up so the new code doesn't become the next legacy

Here is where incremental programs quietly fail. The release reaches 100%, the team moves on, and the toggle, the abstraction and the old implementation stay. A 2021 study in Empirical Software Engineering of 12 Python projects and 61 practitioners found that 75% of toggle components were removed within 49 weeks, so a quarter lived longer and some were never removed. Uber built a tool, Piranha, just to delete code behind stale flags, and reports removing around two thousand of them.

The cost is not only clutter. The authors of Patterns of Legacy Displacement call these seams transitional architecture and are blunt: "you will have to invest in work that will be thrown away," and it must be removed. The extreme case is Knight Capital. In 2012, new trading code reused a flag that had once activated a long-retired function. One of eight servers didn't get the new code, so when the flag was set the old function ran, and the firm lost more than $460 million in about 45 minutes, according to the SEC's order. The mechanism is the same at any scale.

So make cleanup a rule rather than an intention:

Where Fabrica stands

The CLEAR method applies this pattern with the evidence gathered first. Capture instruments the running system, including jobs and queries, so a module's callers are found by observation rather than estimate, and Lock turns what is observed into a System Specification that owners approve. Engineer rebuilds the scoped workflow behind the seam, Attest tests both implementations against the same specification before any customer is moved, and Release moves cohorts along a 1%, 10%, 50%, 100% path with acceptance criteria at each step and rollback as a routing or toggle change. The detail is at how it works and on the engineering page.

The scaffolding comes down

Branch by abstraction and feature toggles let you replace the parts of a legacy system no proxy can reach, one caller and one cohort at a time, with the current system serving until the evidence supports a move. What separates the programs that end well from the rest is rarely the seam or the toggle. It's whether someone was accountable for taking them down. Plan the removal when you plan the release, and treat a toggle still live a year after 100% as the defect it is.

Questions to ask before you cut the first seam

Keep reading

Related

Which priority is your system holding up?

Tell us about it on a 30-minute call. We'll suggest a first scope and what a technical review would need to confirm.

Book a 30-minute call