Resources · Patterns

The strangler fig pattern: replacing a legacy system one workflow at a time

In short

The strangler fig pattern replaces a legacy system one workflow at a time behind a routing layer, while the current system keeps serving everything else. It works because each workflow is captured, specified, rebuilt and proven before traffic moves, and rollback is a routing change. It breaks down at shared tables, background jobs and batch reports unless each gets its own plan.

Abstract illustration: a vine wrapping a grey column, new growth in teal

You run a core system that has been in production for ten or fifteen years. It still does the job, but every change takes longer than it should, and someone has proposed replacing it at least once. The proposal usually comes with a big-bang cutover date, and you've seen how those dates behave.

The strangler fig pattern is the standard alternative. It replaces a legacy system one workflow at a time, behind a routing layer, while the current system keeps serving everything you haven't moved. It's twenty years old, well understood, and still what we'd pick for most systems that matter.

The takeaway: the pattern works because of the routing layer and the discipline of proving each workflow before you move it. It breaks down in three places web requests never touch: shared tables, background jobs and batch reports. Plan for those on day one and the pattern holds; ignore them and the migration stalls.

Where the pattern comes from

Martin Fowler published the idea on June 29, 2004. Strangler figs seed in a host tree's branches, root, and eventually stand on their own while the host dies inside them. A replacement system, he argued, could grow the same way around the old one instead of arriving in one cutover.

His August 22, 2024 revision is more useful to a buyer. It says plainly why the simple plan fails: "we've seen this simple-sounding plan go down in flames most of the time. Replacing a serious IT system takes a long time, and the users can't wait for new features. Replacements seem easy to specify, but often it's hard to figure out the details of existing behavior."

The market data points the same way. In a 2022 Wakefield Research survey for vFunction of 250 US software and architecture leaders at large companies, 79% said at least one modernization effort had failed, with the typical effort costing about $1.5 million and running 16 months. It's a vendor-sponsored survey, so treat the exact figure with care. The direction matches what we've seen over 13 years and 200-plus projects.

The routing layer that makes it possible

The mechanical heart of the pattern is a facade in front of the current system: in a web application, a reverse proxy or API gateway. On day one it sends every request to the legacy application; customers see no change. As each workflow is rebuilt, the proxy sends that workflow's requests to the new implementation and everything else to the old one.

Three things follow from putting the routing layer in first.

There's a cost. AWS's guidance on the pattern notes that the proxy can become a single point of failure, and that calls inside the monolith to a moved workflow need an anti-corruption layer. Fowler's colleagues call this transitional architecture: code you build knowing you'll throw it away. People balk at that, but it's a small fraction of the total and it buys you the rollback.

One workflow at a time: the cycle that fits the pattern

The routing layer tells you where to cut, not how to cut safely. For that, we run five steps for every workflow, in order, and don't start the next until the previous one is released.

  1. Capture. Observe what the running system does for this workflow: requests, queries, background jobs, errors, reports, and the rare month-end paths. Anything not yet observed is recorded as a gap, not assumed.
  2. Specify. Turn the evidence into a written specification the business owners approve: rules, inputs and outputs, data relationships, exceptions, acceptance criteria. Lock the version; later changes create a new one.
  3. Rebuild. Build the workflow on the new stack against the locked specification, alongside the current system, with no customer traffic.
  4. Prove. Run both implementations against the same inputs and compare outcomes, data, reports, permissions and performance. Thoughtworks calls this parallel run with reconciliation: the legacy answer is returned, and differences are resolved before anyone is moved.
  5. Release. Move traffic by cohort, gated on evidence at each step, with acceptance criteria and a tested rollback.

The order matters. Rebuild first and specify afterwards, and you prove only that the new system matches whatever you built.

A publisher's subscription system is a fair example. The first workflow was renewals, because a pricing change was stuck behind it. Capture surfaced a nightly job that extended grace periods for a product line nobody had sold in six years, and a report the finance team adjusted by hand each month. Neither sat in the code the engineers called "renewals." Both went into the specification before anyone wrote new code.

Where the pattern breaks down, and what to do about each

The routing layer sees HTTP requests. A real legacy system does a great deal that never arrives as a request, and that's where migrations stall.

Shared tables

The classic failure is a new service that reads and writes the same tables the monolith still uses. In production, two codebases enforce two versions of the same rule on one table, and nobody owns the data. AWS calls synchronizing two stores "a tactical solution," a polite way of saying it has a shelf life.

What to do: decide, per table, which side owns writes at each stage, and move ownership deliberately. Shopify's engineering team documented the sequence when extracting settings from a 3,000-line Rails model: define the new interface, switch readers to it, create the new store, write to both in a transaction, backfill history, switch reads, then stop writing to the old store and delete it. Seven steps, each reversible. Where the legacy side still needs the data, the new system writes it back in the old shape, a legacy mimic. Data cutover is its own workstream, with its own reconciliation and rollback.

Background jobs

Cron jobs, queue workers and scheduled tasks don't pass through the proxy. If renewals move but the nightly job that charges the cards still runs in the old system, you've moved half a workflow, and the half you can't see has the side effects.

What to do: treat the point where jobs are enqueued as a seam of its own. The event interception pattern covers this: intercept the message, queue entry or trigger and route it the way the proxy routes requests. Two rules we don't bend. At any moment exactly one system owns each job; two schedulers running the same job is how customers get charged twice. And while a job is being proven in parallel, its side effects are captured, not executed: emails, payments and webhooks go to a sink until release.

Batch reports

Month-end close, the board pack, the regulator's file: Fowler's colleagues call these a critical aggregator, and they're why many migrations leave reporting until last, feeding old reports from new systems for years. They're blunt about it: legacy reports "contain undiscovered issues and bugs," so "the new outputs rarely, if ever, match the existing ones."

What to do: map every report's upstream sources during capture, not after. Decide early whether to rebuild the report on the new data or keep feeding the old one from a mimic. Either way, reconcile totals against the current system for a defined period, with the owners ruling on which system is right. "They're off by $400 and we don't know why" is a finding to resolve before release.

Where Fabrica stands

Fabrica's CLEAR method is this cycle with the names fixed: Capture, Lock, Engineer, Attest, Release. The routing layer goes in first, the rebuilt workflow runs beside the current system with its own data and its side effects held back, and traffic moves by cohort only after the proof report supports it (see how it works). The three hard cases above are stated as known limits: jobs, queues and batch reports get their own handling scoped in the assessment, and data cutover runs as a separate workstream.

Start where the business is waiting

Fowler is careful to say the pattern "doesn't make the exercise easy." It makes replacement survivable: each step small enough to prove, each release small enough to reverse, and "both investment and returns occur gradually and visibly."

Choose the first workflow by cost of delay, not technical convenience: the one holding back a priority the business has already agreed on. Insist that capture and a locked specification come before any new code, and that shared tables, jobs and reports are in the plan from the first week. If a proposal can't say how it handles those three, it isn't a strangler fig plan yet. It's a big-bang rewrite with a longer timeline.

Questions to ask before you commit:

Keep reading

Related

Which priority is your system holding up?

Tell us about it on a 30-minute call. We'll suggest a first scope and what a technical review would need to confirm.

Book a 30-minute call