Most rewrites of a working system fail the same way. The new code passes every test the team thought of, goes live, and then meets the rules nobody wrote down: the discount that only applies on the last day of the month, the customer whose IDs start with a zero, the rounding the old system did "wrong" for fifteen years that finance now depends on.
You can't write tests for rules you don't know exist. But you can make the old system tell you about them. Run the new code alongside the old one on real traffic, compare every output, and treat each difference as either a bug in the new code or an undocumented rule in the old one.
This is the pattern we use when we replace parts of a legacy system. It works with the strangler fig approach Martin Fowler described, where a routing layer sits in front of the old system and parts move to new code one at a time.
The shape of it
┌──────────────► legacy ──► response to the user
request ──► router
└──(async copy)─► new code ──► compared, then thrown away
The user only ever gets the legacy response. The new code runs "in the shadow": same input, output recorded and compared, never returned. Nothing changes for users until the outputs match.
1. A shadow call that can't hurt the real request
// shadow.ts
import { compare } from "./compare";
export async function withShadow<T>(
name: string,
input: unknown,
legacy: () => Promise<T>,
candidate: () => Promise<T>,
): Promise<T> {
const result = await legacy(); // the user always gets this
// Fire and forget: the shadow must never slow down or break the request.
setImmediate(async () => {
const started = Date.now();
try {
const shadow = await withTimeout(candidate(), 2000);
const diff = compare(result, shadow);
await recordRun({ name, input, match: diff.length === 0, diff, ms: Date.now() - started });
} catch (err) {
await recordRun({ name, input, match: false, error: String(err), ms: Date.now() - started });
}
});
return result;
}
Three rules keep this safe:
-
The user path never waits for the shadow. If the new code hangs, the timeout and
setImmediatekeep it out of the request. - The shadow must not write. Point the new code at a read-only replica, or give it a "dry run" flag that skips writes, emails, payments and webhooks. A shadow that sends a second invoice email is worse than no shadow.
- Sample if traffic is heavy. Shadowing 5 to 10% of requests is usually enough to find the rare cases within a few weeks.
2. Compare meaningfully, not byte for byte
A raw comparison drowns you in noise: timestamps, generated IDs, key order, 2.50 against 2.5. Normalise before comparing, and be explicit about what you ignore.
// compare.ts
import { diff as deepDiff } from "deep-diff";
const IGNORE = new Set(["generatedAt", "requestId", "traceId"]);
function normalise(value: any): any {
if (Array.isArray(value)) return value.map(normalise);
if (value && typeof value === "object") {
return Object.fromEntries(
Object.keys(value)
.filter((k) => !IGNORE.has(k))
.sort()
.map((k) => [k, normalise(value[k])]),
);
}
if (typeof value === "number") return Math.round(value * 100) / 100; // money to 2 dp
if (typeof value === "string") return value.trim();
return value;
}
export function compare(a: unknown, b: unknown) {
return deepDiff(normalise(a), normalise(b)) ?? [];
}
Keep the ignore list short and reviewed. Every field on it is a field you've decided not to check, so each one deserves a comment saying why.
3. Turn the differences into a worklist
Store each run with its input, the diff and a match flag. Then group the mismatches by the path that differs:
SELECT name,
diff_path,
COUNT(*) AS runs,
MIN(input::text) AS example_input
FROM shadow_runs
WHERE match = false
AND created_at > now() - interval '7 days'
GROUP BY name, diff_path
ORDER BY runs DESC;
Each row is a question for the team. Is the new code wrong, or is the old system doing something nobody wrote down? Either answer is valuable. In our experience the second kind is the reason the shadow phase exists: those rules get written down and tested, so the next person doesn't have to find them the hard way.
Decide in advance which system wins when they disagree. Usually the legacy output is treated as correct until someone shows it's a bug. Otherwise every mismatch becomes an argument.
4. Define "done" before you start
Without an exit rule, shadow runs go on forever. Agree the criteria for each slice up front, for example:
- Match rate of at least 99.9% on normalised outputs over the last N weeks, with every remaining mismatch explained.
- Coverage of the calendar. The shadow has run through at least one month-end, and a quarter-end if the slice touches finance. Month-end is where the undocumented rules live.
- Latency of the new code within an agreed margin of the old.
- A rollback you've actually rehearsed, not one that exists only in a runbook.
5. Then move users gradually
When a slice meets its criteria, flip the router for a small group first:
const useNew = flags.isEnabled("orders.new-pricing", {
tenantId: req.user.tenantId,
region: req.user.region,
});
const price = useNew
? await newPricing(order)
: await withShadow("pricing", order, () => legacyPricing(order), () => newPricing(order));
Start with one branch, region or internal team, watch the same metrics, then widen. Keep shadowing the users who are still on the old path, so you go on comparing until the last of them moves over.
Only switch the old code path off once nothing has called it for an agreed period. Then delete it, so nobody routes to it by accident a year later.
Tools that help
- GitHub's Scientist library (Ruby, with ports to other languages) packages this "run both, compare, return the old result" pattern for code-level refactors.
-
Traffic mirroring in proxies and service meshes (for example Envoy's request mirroring, or the
mirrordirective in NGINX) can copy requests at the network layer when the two systems are separate services. You still need your own comparison step. - Feature flags for the gradual switch, so moving a group back takes seconds, not a deployment.
The cost of skipping it
Shadow running adds weeks to a migration, and that's the main objection. The alternative is finding the undocumented rules in production, through customer complaints and a month-end close that doesn't balance.
We wrote up the rest of the approach, from the routing layer to the cutover, in legacy modernization without downtime.
Your turn: what's the strangest undocumented rule a legacy system taught you? Month-end logic is the usual suspect, but there's always something new.
From the team at Redlio Labs, where we modernize and run business systems.
Top comments (0)