Resources · Engineering

AI agents in legacy modernization: what they do well and where people stay

In short

AI agents are good at the volume of modernization work and poor at the decisions that can't be undone cheaply. Give them declared roles and bounded tasks, and keep plan review, merge, findings and budget behind gates a named person owns. Spend caps, audit trails and reviewed pull requests make those gates real. Rare paths, judgment calls and compliance sign-off stay with people.

Abstract illustration: a build board with agent lanes and a human gate

Every modernization proposal you see this year will have AI agents in it, and some will say "fully automated". You run a core system on Rails, Java or .NET that carries years of business rules, and you have probably watched a big-bang rewrite run late. So the useful question isn't whether agents can write code. They can. It's which decisions you keep, and how that is enforced once agents move faster than anyone can read.

Our view, from running agent-driven sprints against legacy systems, is this: agents are good at volume, and people are good at the decisions that can't be undone cheaply. A modernization that works gives agents the first and keeps the second behind gates a named person owns. "Fully automated" describes the labor, not the accountability.

This piece covers five things:

  1. What current research says about AI-written code (a base rate, not a demo)
  2. The roles agents can hold, and what each does well
  3. The four gates people must own: plan review, merge, findings and budget
  4. The controls that make those gates real: spend caps, audit trails and reviewed pull requests
  5. The limits that don't go away: rare paths, judgment calls and compliance sign-off

Let's take them in order.

Start from the evidence, not the demo

The research from 2024 and 2025 points one way: the code arrives faster, and the cost moves to checking it.

None of this says don't use agents. It says build the process around verification, because "almost right" in a billing system is the expensive kind of wrong.

The roles agents can hold

Agents work best when each has a declared role and a bounded task, the way you'd staff a team. The five we use are:

What they share is consistency on volume and on the paths that traffic actually covers. What they lack is the context that never made it into code: why a rule exists, who depends on it, and what happens if it changes. That is where the gates sit.

The four gates people own

Plan review

Before any code is written, a person reads the architect agent's plan and approves it. This is the cheapest point to catch a wrong assumption: a data mapping that drops a column, a boundary that splits a transaction, a rule the agent treated as a bug.

Merge

Agents open pull requests. People merge them. No exceptions for "trivial" changes, because the trivial ones are the ones nobody reads. The reviewer checks the change against written standards and against the requirement it claims to implement.

Findings

Agents notice things outside their task: a dead branch, an unindexed query, a report that double-counts. These are logged as findings for a person to accept or reject, not acted on. "We don't have time to review findings — the deadline is next quarter." You have less time to find out later why month-end totals changed.

Budget

Someone sets the spend, per item and per day, and decides what happens when an item hits its cap. An agent that burns a budget retrying the same failing test has told you something about the task, and a person should hear it.

The controls that make the gates real

A gate is only as good as the mechanism behind it. Three controls do most of it.

Spend caps. Say a work item is capped at $15 and the day at $120. An item that stalls pauses at $15 and is flagged for a person. The day ends at $120 however many items are queued. Nobody discovers a $4,000 overnight bill, and the cap becomes a signal: items that hit it are the ones with hidden complexity.

Audit trails. Every agent action is recorded against its work item: which specification version it worked from, which requirement it implemented, what it ran, what it cost, and who approved the result. When an auditor asks "why does the new system do this?", the answer is a chain of links, not a reconstruction from memory.

Reviewed pull requests in your repository. The code lives where your engineers already work, goes through your CI, and is merged by your people. This keeps a modernization from becoming a black box you rent. It also matches where regulation is heading. Article 14 of the EU AI Act defines human oversight as the ability to understand a system's limits, stay alert to automation bias, override its output and stop it. Those are the properties of a good merge gate.

The limits that don't go away

Three kinds of work stay with people however capable the agents get.

Rare paths. Agents learn from what they observe, and some behavior is rare by design. On a publisher's subscription system we worked on, a code branch handled print subscribers on a grandfathered annual plan. It ran once a year, in a window the capture had not yet covered, and the analyst agent proposed removing it as dead code. The owner who reviewed the plan knew those readers by name. The fix was not a smarter agent. It was recording the path as a gap, extending capture to cover it, and keeping a person at plan review.

Judgment calls. Is the system invoicing at zero when there's no rate card a bug or a rule? The code can't tell you. An agent will pick one answer and implement it confidently. Someone from finance has to decide, and the decision has to be written down as a requirement with their name on it.

Compliance sign-off. In regulated sectors, change management expects a named approver who is not the author, and an auditor wants to see that approval with a date. An agent can prepare the evidence pack. It cannot carry the liability. If a vendor says the approval step is optional, ask who signs the audit response.

Why "fully automated" still needs an approval chain

There is no contradiction in agents doing most of the work while people hold every decision that touches production. That is how a well-run engineering team already operates: the people who write the code don't alone decide it ships. Agents change the ratio of writing to reviewing. They don't remove the review.

So when you evaluate a proposal, look past the automation claim to the chain behind it. Who approves the plan? Who merges? Where do findings go? What stops the spend? Where is the record? If the answers are "the system handles it", you are being asked to carry risk the vendor has declined to hold. If the answers are names, roles and a board you can open, you have something you can govern. Then measure a pilot on one workflow against a baseline you agree, and let the numbers, not the demo, decide whether to continue.

Where Fabrica stands

Fabrica's CLEAR method applies this division of labor directly. Business analyst, architect, developer, QA and devops agents move work items through declared stages on a shared board, while your engineers approve plans, merge pull requests in your repository, accept or reject findings and set the spend caps. No change reaches production without your team's approval, and each release is gated on proof against the running system with a tested rollback. Our hypothesis is that this keeps the speed of agents and the accountability of people; a pilot on one workflow is how we'd test it with you.

Questions to ask an agent-based modernization partner

Keep reading

Related

Which priority is your system holding up?

Tell us about it on a 30-minute call. We'll suggest a first scope and what a technical review would need to confirm.

Book a 30-minute call