DEV Community

Sundar Shyam Jha
Sundar Shyam Jha

Posted on

7 AI tools for legacy modernization in 2026, and what 507 work orders say about picking one

7 AI tools for legacy modernization in 2026, and what 507 work orders say about picking one

Rewriting an existing system needs about four times more written "definition of done" per unit of work than building a new one. I got that from a public dataset of nine AI-assisted projects: eight modernizations averaged 1.5 acceptance criteria per story point, the one greenfield build averaged 0.4. If you are choosing an AI modernization tool this year, that ratio is a better buying question than any benchmark score.

Here is the method, then the seven tools, then how I would actually sequence a migration.

The number, and how I got it

SoftwareForge published a 2026 enterprise benchmark covering 507 work orders and 3,410 story points across nine initiatives, with the acceptance-criteria density listed per stack. Acceptance criteria are the checkable statements attached to a unit of work, things like "the endpoint rejects unauthenticated requests with 401" or "coverage is at or above 90%."

I divided story points by work orders to get the average size of one unit of change, then divided acceptance criteria by that to normalize. Small table, run it yourself:

Project Story pts / WO AC / WO AC per story point
.NET Framework + WCF monolith to microservices 8.45 11.3 1.34
T-SQL to Aurora MySQL 6.46 10.4 1.61
T-SQL to PL/pgSQL 6.79 10.6 1.56
Legacy framework + jQuery to Vue 3 6.41 8.9 1.39
Java Spring in-place hardening (healthcare) 7.05 8.5 1.21
Java servlets to Node/TS + React 4.33 6.6 1.53
JavaScript to typed React SPA 3.55 9.0 2.54
.NET Core 2.1 (EOL) to .NET 8 2.79 5.0 1.79
Greenfield multiplatform + Terraform 7.93 3.0 0.38

Median for the eight migrations: 1.55. The greenfield project: 0.38.

Two honest caveats. There is one greenfield row, so this is a signal and not a law. And these are planning artifacts, not measured outcomes, so the density reflects how much someone thought needed pinning down, not how much turned out to matter. Useful reading, all the same: when you preserve behavior you cannot describe, the plan has to carry the description. When you invent behavior, you get to change your mind.

If you want to check your own backlog against this, the calculation is four lines:

import csv

with open("backlog.csv") as f:  # columns: story_points, acceptance_criteria
    rows = [r for r in csv.DictReader(f)]

sp = sum(float(r["story_points"]) for r in rows)
ac = sum(float(r["acceptance_criteria"]) for r in rows)
print(f"{len(rows)} work orders, {ac/sp:.2f} acceptance criteria per story point")
Enter fullscreen mode Exit fullscreen mode

Under 0.5 on a migration backlog means your agents are guessing at behavior you never wrote down.

Why this is the buying question

Legacy modernization rarely goes wrong at the translation step. A converter turns a PERFORM loop into a for loop in an afternoon. No converter can tell you why that loop skips every third record when an account flag is set, because nobody ever wrote that rule anywhere except inside the loop.

Coding agents have the same blind spot, at higher throughput. In Stripe's rollout of Claude Code to 1,370 engineers, one team moved 10,000 lines of Scala to Java in four days against a ten-engineering-week estimate. Read the same page for the part people skip: their infrastructure lead describes the mental model that made it work, treating the assistant as a capable new engineer who knows every language but has no business context and no idea how things are done here. Speed came from the context they supplied, not from the model.

Without that context, the failures are quiet. DORA's 2025 report found AI adoption correlating with higher throughput and higher instability at the same time: more change failures, more rework. GitClear's 2025 code-quality analysis tracked duplicated blocks rising as AI-assisted commit volume rose. And in Stack Overflow's 2025 survey, the top frustration among developers using AI was output that is "almost right, but not quite," with debugging that near-miss costing more than writing it fresh.

The tool question, then, is not "can it generate Java from COBOL." All seven below can produce plausible target-language code. What separates them is what happens in the hours before generation starts.

The seven

I have grouped these by what they do first, because that is the actual difference.

1. SoftwareForge (Forge)

SoftwareForge

Disclosure up front: this list came out of research I did for SoftwareForge, and the dataset above is theirs. Judge the number on its method, not on my byline.

Forge starts with a repository scan called ForgeScore, an eight-dimension health assessment across security, architecture, performance and AI adaptability, run before any migration plan gets written. The output becomes a Living Specification, a versioned machine-readable document holding business objectives, functional requirements, security obligations and an explicit out-of-scope boundary. Agents then execute Work Orders generated from that spec, each carrying its own acceptance criteria and a link back to the requirement that authorized it.

Forge bets that the expensive discovery work should happen once, deterministically, rather than being repeated by every agent on every task. Their own writeup on modernizing without breaking business logic is more useful than the product pages if you want the reasoning.

Where it fits: regulated portfolios, multi-quarter migrations, anywhere someone will eventually ask which requirement authorized a specific line in production. Where it does not: a two-week framework upgrade. The planning layer is overhead you will not recover on small work.

2. GitHub Copilot app modernization

GitHub Copilot app modernization

Microsoft ships modernization paths for Java and .NET as guided upgrade flows: assess the project, apply the upgrade, fix what breaks, iterate. The .NET modernization overview and the Java upgrade docs show the actual shape, and GitHub's own modernize legacy code tutorial covers the general pattern.

Strongest when the target is well-known and the gap is version drift rather than architecture. Framework and runtime upgrades, dependency remediation, Azure-bound refactors. This path assumes you already know what the system is supposed to do.

3. Amazon Q Developer transformation

Amazon Q Developer transformation

AWS runs transformation as a defined job type rather than a chat: you point it at a Java or .NET codebase and it produces an upgrade plan and applies it, documented in the code transformation guide.

Same trade as the Microsoft option, on the other cloud. If your estate is already on AWS and the work is Java 8 or 11 moving forward, this is the shortest path. This gets narrower than it looks the moment your migration involves a database dialect change, which is exactly where the acceptance-criteria density in that table went highest.

4. IBM watsonx Code Assistant for Z

IBM watsonx Code Assistant for Z

The mainframe case. IBM's product works COBOL through an explicit sequence: understand the application, refactor the business services, then transform to Java, with the option to stop after any stage.

Staging like that is the interesting part. Stopping after "understand" and shipping nothing is a feature, not a limitation, and it is the honest answer for estates where nobody currently knows what the system does.

5. Moderne and OpenRewrite

Moderne and OpenRewrite
The deterministic option. OpenRewrite parses source into a lossless semantic tree and applies recipes, which are code, not prompts. Same input, same output, every run. Moderne is the commercial platform that runs recipes across many repositories at once.

I would put this first for anything mechanical and repeated: framework migrations, API deprecations, dependency bumps across 300 services. Recipes cannot infer intent, and Moderne does not pretend otherwise. When a recipe exists for your change, using an LLM instead is a worse trade.

6. Devin

Devin
An autonomous agent that takes a task and runs it to completion with little supervision. Independence is the whole value and the whole risk here, because a vague goal produces a confident implementation of the wrong thing.

Devin gets much better when it is fed bounded work with explicit in-scope and out-of-scope boundaries. Same claim as the rest of this post, from a different direction: the agent is not the variable, the brief is.

7. Mechanical Orchard

Mechanical Orchard
Behavior-first mainframe replacement. Rather than translating source, the approach observes what the running system actually does and builds a modern implementation to match that observed behavior.

Of everything here this is the slowest and most expensive, and for a 40-year-old system with no surviving documentation and no surviving authors, it is sometimes the only defensible one.

How I would actually sequence this

Nothing above works if you skip the step everyone skips.

Before any tool touches the code, write characterization tests. Michael Feathers named these in Working Effectively with Legacy Code, and the definition matters: a characterization test does not check that code does what it should. It records what the code currently does for real inputs, so every later change is measured against observed behavior rather than assumed intent. Writing one for an undocumented function is usually the fastest way to find a business rule that no design document mentions.

Then the order I would run:

  1. Scan and score, so priority is defensible to a reviewer rather than convenient for whoever is free
  2. Characterize the twenty modules carrying the most transaction volume
  3. Take the deterministic wins first (recipes over prompts wherever a recipe exists)
  4. Scope the rest into work orders small enough that one person can verify the acceptance criteria without re-reading the diff cold
  5. Route each work order to the cheapest model that clears it

And be careful about the dependency layer, because it fails silently. There is no single command that audits a monorepo of independently managed services:

cd services/paymentservice
mvn dependency:tree > dep-tree-payment.txt
cd services/inventoryservice
mvn dependency:tree > dep-tree-inventory.txt
# repeat per service, then cross-reference each output against a CVE database
Enter fullscreen mode Exit fullscreen mode

A single service tree runs several hundred lines. An agent adding a legitimate payment SDK can pull a vulnerable auth library four levels down and nothing in the pull request will say so.

The part the vendor decks leave out

You should expect less speedup than the demos suggest, at least at first.

METR ran a randomized controlled trial with experienced open-source developers on their own repositories and found they took 19% longer with AI tools while believing they had been sped up. Plenty of people quote that result as a permanent verdict, which is not fair to it. METR revisited the design in February 2026 and said plainly that developers are likely faster now, but that their newer data has selection problems severe enough that they will not put a number on it.

Carry both halves of that. Familiarity with the codebase reduces the gain, self-reported speedup is unreliable, and the honest current answer on magnitude is that nobody has a clean measurement.

Which brings it back to the ratio. The gain is not in the generation. Gains come from how much of the system's behavior you managed to write down before generation started. You can measure it on your own backlog this afternoon, and it does not require buying anything.

If your migration backlog is sitting under 0.5 acceptance criteria per story point, no tool on this list will save you.

Top comments (0)