DEV Community

Cover image for A refactor without tests is a rewrite
Hayati
Hayati

Posted on Originally published at hayatibis.dev

A refactor without tests is a rewrite

Originally published on hayatibis.dev, where the figures move and the toys are interactive.

“I’m refactoring it” is one of the most stretched sentences in software. It gets used for renaming a variable, and for a service that is down for a week while someone cleans it up.

Martin Fowler, who wrote the book on it, puts it more precisely. A refactoring is “a change made to the internal structure of software to make it easier to understand and cheaper to modify without changing its observable behavior.”

The interesting part is the end of that sentence: without changing its observable behavior. It raises a question people tend to skip: how do you know?

How do you know nothing changed?

You can read the diff, or try a few things in the app (these days we often don’t even read PR diffs: an AI agent writes the code, and another one reviews it :) Neither holds up once the change is bigger than a few lines.

Tests are the simple answer. They pin down what the code does today, and they tell you within seconds when a step moves it. Without them, “the behaviour didn’t change” is a wish rather than a fact. A change you can only hope is safe carries the risk of a rewrite, whatever the ticket calls it.

Fowler has a one-line test for this: “If somebody talks about a system being broken for a couple of days while they are refactoring, you can be pretty sure they are not refactoring.”

Fowler would call the untested kind restructuring rather than a rewrite. My rule is stricter: if I can’t show the behaviour stayed put, I plan the change with the care I’d give a rewrite.

Small steps, each one checked

Refactoring, done his way, is a chain of small moves with names: Rename Variable, Extract Function, Move Function. After each one you run the tests.

These are the moves we worked hard to get used to, and racked our brains over, back before AI. Now, building with Claude and Codex, we sometimes have to steer the agents down the same road. The most capable models already know all of this, and the way they are trained, to keep running the code and fixing it until the tests pass,[1] lets them do it in the natural flow of work. Still, we need to take care to hand them our own way of working.

Flip the switch below and watch the same five steps with and without tests.

Fig. 1 · Five small steps. With tests, the step that changes behaviour fails at once and gets undone. Without them, the bug shows up weeks later and nobody knows which step it came from.

Fig. 1 · Five small steps. With tests, the step that changes behaviour fails at once and gets undone. Without them, the bug shows up weeks later and nobody knows which step it came from. (interactive version)

With tests, a broken step is always the last step, so you know exactly what to undo. Without them, all five ship together. When the bug arrives, the question is no longer “what did I just do?” but “what did we do in October?”

The legacy code trap

Michael Feathers put it bluntly in Working Effectively with Legacy Code: legacy code is simply code without tests. Age has nothing to do with it. Last week’s code without tests is legacy code.

And it comes with a trap. To change the code safely you need tests. To get tests around it you often have to change the code first: a constructor that opens a database connection, or a function that reads the clock. Feathers’ way out is a loop where the change you actually came for is the last step.

Fig. 2 · Feathers' legacy code change algorithm. The first edits only make the code testable; the change you came for comes last.

Fig. 2 · Feathers' legacy code change algorithm. The first edits only make the code testable; the change you came for comes last. (interactive version)

Step 3 is where the careful, tiny edits go: pass the clock in as a parameter, pull the database behind an interface. Feathers calls the places where you can swap behaviour without editing the code there seams.

Step 4 has a trick I like. You don’t write tests for what the code should do; you write characterization tests for what it does. Write an assertion you know is wrong, run it, and let the failure tell you the real answer. Then put the real answer in the test.

// A characterization test: it records what Price does today, not what it should do.
// If today's answer looks wrong, pin it anyway and fix it in a separate change.
func TestPrice_CurrentBehaviour(t *testing.T) {
    for cents, want := range map[int]string{
        0:      "0.00",
        199:    "1.99",
        -5:     "-0.05",
        100000: "1,000.00",
    } {
        if got := Price(cents); got != want {
            t.Errorf("Price(%d) = %q, want %q", cents, got, want)
        }
    }
}
Enter fullscreen mode Exit fullscreen mode

Now the refactor has a net under it.

Refactor or rewrite?

Fig. 3 · Three changes. By my rule, which one is a refactor?

Fig. 3 · Three changes. By my rule, which one is a refactor? (interactive version)

When it really is a rewrite

Rewrites aren’t a sin. Sometimes the old code has to go. They just need a different plan, and the first step of that plan is to define the rewrite in a much clearer frame and say its name louder.

A rewrite you’ve named can be planned like one: run the old and new code side by side and compare what they return, move traffic behind a flag, keep the way back open. Fowler’s strangler fig is the patient version of this: grow the new code around the old, a piece at a time, until the old part can be switched off.

Next time someone says “I’m refactoring it”, ask the question this whole essay hangs on: how will you know nothing changed? If the answer is a test suite, let them refactor as much as they like, in small steps. If the answer is “I’ll click around”, it’s a rewrite with a nicer name, and it deserves a rewrite’s plan.


Stands on


Notes

  1. OpenAI says codex-1, the model behind Codex, was trained with reinforcement learning on real-world coding tasks and “iteratively runs tests until passing results are achieved.” Anthropic’s Claude Code guide describes the same loop: Claude “does the work, runs the check, reads the result, and iterates until the check passes.”

If you'd like more of these, you can support the writing on Patreon.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

Official Platform Update

Security protocols have been updated for all developer accounts.

  • tr.ee/dev-to