Resources · Engineering

Rollback is a routing change: designing releases you can undo

In short

A release you can undo is designed, not hoped for. Put a routing layer in front of the current system, move production by cohort starting with a beta opt-in, write acceptance criteria for each step and rehearse the rollback before go-live. The hard part is data: expand/contract for schema, one owner per write, and a ledger of side effects already sent.

Abstract illustration: a routing switch with a lever to return traffic

If you have sat through a go-live weekend for a core system, you know the feeling. The new version goes in overnight, the old one is switched off, and from Monday the only way back is a restore from backup and an awkward conversation with the board. Everyone has agreed the plan is "low risk". Everyone is quietly hoping.

Rollback is usually the last slide in the release deck. It should be the first design decision. When the way back is cheap, you can move a small cohort, watch it, and carry on or step back without drama. When it is expensive, you are betting the business on one night.

The takeaway: design the release so that undoing it is a routing change, not a restore. In our experience that comes down to five things, and the difficult ones are about data, not traffic.

  1. A routing layer that owns the switch (one place decides which system serves which request)
  2. Cohorts small enough to learn from (a beta opt-in group first, then percentages)
  3. Acceptance criteria for each step (written before the step, not argued after it)
  4. A data plan that keeps both systems valid (schema, write ownership and side effects)
  5. A rollback you have already run, and a log of every transition

Let's take them in turn.

1. A routing layer that owns the switch

None of this is new. Martin Fowler described blue-green deployment in 2010: two production environments with a router in front. Test the new version in the idle one, switch the router, switch it back if anything goes wrong. Danilo Sato's canary release (2014) routes a small subset of users first and widens as confidence grows; rollback is rerouting them to the old version.

For a legacy renewal the same shape applies. A reverse proxy sits in front of the current system and, on day one, sends it everything. The rebuilt workflow receives no customer traffic until the proof says it should. When a cohort moves, a routing rule changes. When it comes back, the rule changes again. No redeploy, no restore.

DORA's 2024 report has the best-performing teams recovering from a failed deployment in under an hour, which no restore from backup achieves. The counter-example is Knight Capital in 2012: new order-routing code deployed, one server missed, and in 45 minutes the router sent more than 4 million orders and lost the firm more than $460 million (SEC, 2013). Nobody had designed a way to turn it off.

2. Cohorts small enough to learn from

A random percentage of traffic is the usual canary, and for a business workflow it is often the wrong unit. An invoice half-processed by each system teaches you nothing and creates a reconciliation problem. Route by something that owns a complete business cycle (a customer, a tenant, a region) and keep each cohort on one system for the whole cycle.

The lowest-risk first cohort is one that chose to be there. A beta opt-in puts a small banner in front of customers: a new version of this page is available, try it, switch back any time. Customers who opt in are forgiving, they tell you what broke, and the switch back is their decision. You learn how the rebuild behaves with real data before moving anyone who did not ask.

After opt-in, an illustrative path is 1%, 10%, 50% and 100%, each step gated on evidence from the one before. The numbers matter less than the gates. In July 2024 a CrowdStrike content update went to every sensor at once and, by Microsoft's estimate, 8.5 million Windows devices crashed. CrowdStrike's post-incident review committed to a staggered rollout starting with a canary: the cohort model, adopted after the fact.

3. Acceptance criteria for each step

"We'll watch the dashboards." Which ones, for how long, and who may call it back at two in the morning without a meeting? Acceptance criteria answer those questions before the step, while nobody is tired or defensive. For each one, agree:

4. A data plan that keeps both systems valid

This is where "rollback is a routing change" stops being trivially true. Routing is easy to undo. Data is not. Three problems come up in almost every renewal.

Data written by the new system

If the rebuilt workflow writes to its own store and you route a cohort back, those records vanish from the current system's view. Fowler calls this "the issue of dealing with missed transactions".

The discipline is one owner per write, at every moment, with the other side kept as a copy. Stripe's engineers described the pattern they use for online migrations: dual-write to old and new, switch reads while comparing results against the old store, switch writes, and only then remove the old data. A reconciliation job proves both stores agree, and any difference is a defect with an owner. Rolling back then means the current system picks up records it already has.

Schema changes

A new version that drops a column the old version reads has removed your rollback. The expand/contract pattern (Sato's "parallel change") splits every schema change into three phases: expand so both versions work, migrate consumers to the new shape, and contract once nothing depends on the old one. Our rule: contract never happens while a step is inside its watch period. If you might roll back, the old shape stays.

Side effects already sent

An email, a payment instruction, a shipping label or a call to a partner's API cannot be un-sent. Rolling back the routing does not roll back the world. Two practices help. Idempotency: every side effect carries a key derived from the business event, so if the current system sends it again after rollback, the receiver sees a duplicate. And a side-effect ledger: what was sent, by which system, for which event, so you know what must not be repeated. Where neither works, decide before the step what the compensating action is and who performs it.

A German fintech client learned this the uncomfortable way, before these controls were in place. A rebuilt onboarding workflow was rolled back after a limit check differed from the current system. The rebuild had written verification status to a new table only, so the current system blocked a few dozen accounts as unverified. The routing change took a minute. Reconciling the accounts by hand took most of a day.

5. A rollback you have already run, and a log

A rollback plan that has never been executed is a hypothesis. Run it before go-live: route the beta cohort to the new version, run a business cycle, route it back, reconcile. Measure how long it took and which records needed attention. As Fowler notes, this is the same mechanism as a hot standby, so you test recovery on every release instead of once a year.

Then log every transition. The log is what you show an auditor, and what the next workflow learns from. For each step, record:

Without those six lines, a rollback at month-end turns into archaeology.

Where Fabrica stands

Release is the fifth stage of Fabrica's CLEAR method, and it follows the shape above. A routing layer in front of the current system moves traffic to the rebuilt workflow by cohort, starting with a beta opt-in, with each step gated on evidence and acceptance criteria your team signs off. Rollback is a routing change tested before go-live, data cutover is its own workstream, and a shift plan and log record each transition. The stages are on how it works; what your engineers can check is on engineering.

Start with the way back

Most release plans describe the way forward in detail and the way back in a sentence. Reverse the emphasis. If you can say, for each step, how a cohort comes back and what happens to the data it created meanwhile, the forward path almost writes itself. If you cannot, the release is a bet. Routing is the easy part. The work that makes it honest is in the cohorts, the criteria, the data plan and the rehearsal. Do that work first, and a go-live stops being a weekend.

Questions to ask before any release step

Keep reading

Related

Which priority is your system holding up?

Tell us about it on a 30-minute call. We'll suggest a first scope and what a technical review would need to confirm.

Book a 30-minute call