Resources · Decision guide
In short
Twelve questions, grouped for finance, risk and engineering, that show whether a modernization vendor has renewed a live system before: baseline, pilot terms, reporting, exit, rollback, access, data, documentation, proof, freeze, code ownership and the willingness to say don't rebuild. The red flags are a missing baseline, no rollback mechanism, documentation generated from code alone and any promise of zero risk.

You have two or three proposals on the table for the system that is holding the business back. Each promises a modern stack, a short timeline and low risk. Side by side, you can't tell which vendor has done this before and which is about to learn on your production system.
The proposals won't tell you. The answers to twelve specific questions will. A vendor who has renewed a live system has already thought about baselines, proof, rollback and exit, and will answer concretely. One who hasn't will answer with adjectives. Here are the questions we'd put to any vendor, ourselves included, grouped by who should ask them.
For finance
For risk and IT
For engineering
Let's see what a good answer to each one sounds like.
Every proposal claims to be faster and cheaper than your current approach. Faster than what? Ask for the figures the comparison rests on: cost per change, time from request to production, internal hours on upkeep. A credible vendor wants you to measure the baseline yourself, before work starts, so the result can't be argued with later. McKinsey and Oxford's study of 5,400 large IT projects found they ran 45% over budget on average and delivered 56% less value than predicted. Without a baseline, you won't know whether you're one of them.
A pilot should have a fee agreed before work starts and a scope narrow enough to finish. The revealing half is the second one. If you stop, do you walk away with a specification, evidence and a recommendation you can use with anyone, or with a half-built module only this vendor can continue?
Ask to see the report format before you sign, not a description of it. You want what moved, what it cost, what is waiting on a decision from your side, and what is still unknown. Ask whether spend is capped per item and per period. A monthly status deck is a weak answer. A board you can open at any time is a strong one.
Exit terms are easiest to agree when nobody expects to use them. Ask what is handed over, in what state, and whether anything you need runs on infrastructure you don't control. The UK Government Digital Service's guidance on technical lock-in says it plainly: retain ownership of intellectual property and data, and put the estimated cost and time to exit in the business case.
This question separates vendors who have moved live traffic from those who haven't. Listen for mechanics, not reassurance. Is rollback a routing change or a redeploy? Was it tested before it was needed? Does the current system keep running until each workflow is fully moved?
Fowler's strangler fig pattern works because the old system stays in service while the new one grows around it. DORA's research adds that smaller changes are easier to recover from. One large cutover means one large recovery.
For a first conversation, none. For an assessment, read-only access to code, schema and analytics, agreed with your engineers. Ask how payloads are redacted and what customers see if collection fails. Be wary of anyone who wants production write access in week one.
Workflow logic gets the attention. Data cutover is where migrations hurt. Ask whether data transition is its own workstream, with mappings, backfill, write ownership and reconciliation defined before live writes route to the new implementation. Then ask who reconciles the month-end totals against the current system.
Most proposals include "full documentation of the legacy system". Ask how it is produced. Documentation generated from code alone describes what the code does, not why, and misses rules that live in data, configuration and people's memory. Spolsky made the point in 2000: old code is full of bug fixes that each took weeks of real use to find, and none of them are labelled.
A publisher we worked with had a subscription renewal rule that existed only in a background job and a support agent's head. It surfaced when observed behavior and the code were put in front of the owner together. Ask how observation, code and owner review are combined, and how pages stay current as the system changes.
Feathers defined legacy code as code without tests, so the vendor can't lean on your test suite. Ask what the specification is, who approves it, and how both implementations are tested against it. Replay or shadow of representative workloads, with side effects controlled, is the answer you want: compare outcomes, data, reports, permissions and performance, and resolve differences before any customer moves. "We'll test it thoroughly" is not proof. A report listing each difference found and how it was closed is.
The default answer should be no. Parallel builds and incremental replacement work without a whole-system freeze. If a scope needs one, usually around data cutover, the vendor should say what it covers and for how long. A proposal that quietly assumes six frozen months is asking the business to stop.
In your repository, as pull requests your engineers review. If AI agents do the building, ask what humans approve: plans, merges, configuration, budget. When an agent notices something outside its task, it should be logged for a person to decide, not acted on.
Not every system needs rebuilding; upgrading in place or refactoring a narrow part is often the better call. A vendor who can't describe when they'd recommend against their own service is selling capacity, not advice. Make "don't rebuild" a permitted outcome of the assessment.
In our experience, the proposals that hold up share a shape. One bounded workflow, chosen because it sits behind a business priority, not because it is easy. A fee agreed up front. Acceptance criteria written with your owners. A baseline you measure. A rollback tested before go-live. And a decision point at the end where "stop" is a legitimate answer.
Notice, too, what the vendor declines. A credible one should say no to:
"But we need the whole thing done by next year." You may. A vendor who says yes before measuring anything is telling you what you want to hear.
Fabrica's CLEAR method is built around these questions. Capture observes the running system with read-only access, Lock turns the evidence into a System Specification your owners approve, Engineer rebuilds one workflow as reviewed pull requests in your repository, Attest tests both implementations against the same specification, and Release moves traffic by cohort with a rollback that is a routing change. The assessment measures a pilot against a baseline you agree and may recommend not rebuilding; the finance page lists what we won't claim.
You don't need to be an engineer to ask these questions, and you don't need a vendor to score twelve out of twelve. You need to know which answers were specific and which were adjectives. Put the specific ones in the contract: the baseline, the pilot scope, the report format, the exit terms. Then judge the vendor on whether the pilot's measured results match what they said in the room. A good one would rather be compared on evidence than on slides.
Tell us about it on a 30-minute call. We'll suggest a first scope and what a technical review would need to confirm.