DEV Community

Dakota Huang
Dakota Huang

Posted on

One Helper Moves Only After the Status Table Stays Green

One helper moves only after the status table stays green. A messy classifier mixes folding, alias lookup, and logging. Those three jobs should not move in one patch.

This workflow freezes observed outputs, then extracts one string helper. The helper must not change any pinned label. A later patch can fix quirks after the freeze holds.

The code below is an unexecuted teaching fixture for a scratch repo. It is not a claim about any live order system. Copy it, run it, and keep the results next to your diff.

The conclusion in one rule

Preserve today's outputs before you improve today's structure. Structure work is safe only when the table stays identical. Behavior work belongs in a separate, reviewed change.

That rule fits messy repositories with unclear ownership. It fits less well when the ticket demands a new label now. Read the skip section before you adopt the steps.

What the fixture actually does

The function classify_order accepts raw text and returns one label. A None input becomes an empty string before lookup. A known typo, the word payed, still returns paid.

Spaced input such as pa id returns review. That row is a quirk, not a target design. Empty input returns unknown, and other text returns pending.

Each call also appends a log line to a caller-supplied sink. The line stores the original text and the label. Tests assert both the label and that line.

The integer input stays in the table because callers may pass numbers. Removing that coercion is a behavior change, even if it looks cleaner. Leave it until a separate patch updates the expected row.

Fixture code

def classify_order(raw, sink):
    text = "" if raw is None else str(raw)
    folded = text.strip().lower()
    aliases = {
        "ok": "paid",
        "paid": "paid",
        "payed": "paid",
        "cancel": "cancelled",
        "canceled": "cancelled",
        "cancelled": "cancelled",
        "": "unknown",
    }
    if folded in aliases:
        label = aliases[folded]
    elif " " in folded and folded.replace(" ", "") in aliases:
        label = "review"
    else:
        label = "pending"
    sink.append(f"{text}|{label}")
    return label
Enter fullscreen mode Exit fullscreen mode

Place the function in order_status.py before you write tests. Keep the alias literals inline for this first freeze. Moving the dictionary is a later cut, not this one.

Step 1. Execute rows before you write expectations

Run the function for a fixed input list and print both results. Record those prints as the golden table for this extract. Do not invent a row you have not executed.

Use this runner in the same scratch directory as the fixture.

from order_status import classify_order

def main():
    samples = [None, "", "  ", "paid", " payed ", "pa id", "CANCEL", "nope", 12]
    for raw in samples:
        sink = []
        label = classify_order(raw, sink)
        print(repr(raw), label, sink[0])

if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode
python status_probe.py
Enter fullscreen mode Exit fullscreen mode

The probe is part of the method, not a benchmark harness. Save its printed output beside the test file. If a row surprises you, keep it until a later behavior patch.

Step 2. Turn the probe output into a table

input     label       log line           note
None      unknown     |unknown           None becomes empty text
""        unknown     |unknown           empty stays unknown
"  "      unknown     "  "|unknown       strip folds spaces to empty
paid      paid        paid|paid          direct alias
" payed " paid        " payed "|paid     typo and padding kept
pa id     review      pa id|review       internal space quirk
CANCEL    cancelled   CANCEL|cancelled   case fold only
nope      pending     nope|pending       fallback
12        pending     12|pending         non-string coerced
Enter fullscreen mode Exit fullscreen mode

These nine rows are the contract for the next extract. The notes explain history, but assertions use only label and log line. A reviewer should be able to recompute every cell from the fixture.

Step 3. Lock the table in one test

import pytest
from order_status import classify_order

CASES = [
    (None, "unknown", "|unknown"),
    ("", "unknown", "|unknown"),
    ("  ", "unknown", "  |unknown"),
    ("paid", "paid", "paid|paid"),
    (" payed ", "paid", " payed |paid"),
    ("pa id", "review", "pa id|review"),
    ("CANCEL", "cancelled", "CANCEL|cancelled"),
    ("nope", "pending", "nope|pending"),
    (12, "pending", "12|pending"),
]

@pytest.mark.parametrize("raw,label,line", CASES)
def test_status_rows_stay_fixed(raw, label, line):
    sink = []
    assert classify_order(raw, sink) == label
    assert sink == [line]
Enter fullscreen mode Exit fullscreen mode

The test module has one job: freeze current behavior. It does not document the ideal domain language. Add a comment that payed and review are preserved on purpose.

Install pytest in that environment before you record the baseline. A missing dependency is a harness failure, not a product failure. Fix the environment, then run the suite again.

Step 4. Record a baseline command

Run the same command you will use after the edit.

python -m pytest tests/test_order_status_rows.py -q
Enter fullscreen mode Exit fullscreen mode

Expect nine passed tests before any production edit. Write that count in the pull request notes. A failed baseline means the harness is wrong, so fix the harness first.

Do not switch flags between the baseline run and the later run. A quieter flag or a new plugin can hide a real miss. Keep the working directory and Python environment fixed too.

Step 5. Score candidate cuts with a local rubric

Use a three-point rubric you can apply by reading the diff. Award one point for each extra function touched. Award one point if a golden label could change.

Award one point if logging text could change. These scores are a local rubric, not a measured field study. A lower score means a narrower change for this gate.

Candidate Touch Label risk Log risk Score Gate
Extract fold_status only 1 0 0 1 allow
Extract alias dict only 1 1 0 2 defer
Extract fold and alias together 2 1 0 3 reject now
Delete the payed alias 1 1 1 3 reject
Replace review with paid 1 1 1 3 reject

Only a score of one may enter the current patch. The allowed cut creates fold_status and leaves lookup in the caller. That split is mechanical and easy to review.

The deferred cut moves policy, so it waits for its own tests. Deleting payed would change a pinned label and a log line. That deletion is a behavior patch, so the gate rejects it.

Step 6. Apply the single allowed extract

def fold_status(raw):
    text = "" if raw is None else str(raw)
    return text, text.strip().lower()

def classify_order(raw, sink):
    text, folded = fold_status(raw)
    aliases = {
        "ok": "paid",
        "paid": "paid",
        "payed": "paid",
        "cancel": "cancelled",
        "canceled": "cancelled",
        "cancelled": "cancelled",
        "": "unknown",
    }
    if folded in aliases:
        label = aliases[folded]
    elif " " in folded and folded.replace(" ", "") in aliases:
        label = "review"
    else:
        label = "pending"
    sink.append(f"{text}|{label}")
    return label
Enter fullscreen mode Exit fullscreen mode

The helper returns original text and the folded key. The caller still owns every policy branch in this file. The log line still uses the original text, not the folded key.

Step 7. Re-run the baseline and stop on any miss

python -m pytest tests/test_order_status_rows.py -q
Enter fullscreen mode Exit fullscreen mode

Require nine passed tests, matching the baseline note. If any row fails, revert the extract immediately. Do not edit an expected label to recover a green bar.

A green rerun means this structural move is safe to review. It does not mean the review quirk is acceptable forever. Schedule that correction as a behavior change with new tests.

Step 8. Optional draft from a free model

Disclosure: This article was prepared as part of MonkeyCode's product outreach.

MonkeyCode offers free model access and a free server option. Those two facts came from the operator for this draft. This article states no model names, quotas, hardware, or expiry.

Use the free model only as a patch drafter for the allowed cut. Paste the fixture, the test, and the rubric row you chose. Ask for fold_status and the matching caller change, nothing else.

Task: extract fold_status from classify_order.
Constraints: do not change labels, aliases, or log text.
Constraints: do not edit tests to make them pass.
Return: a unified diff for order_status.py only.
Enter fullscreen mode Exit fullscreen mode

The free server option can execute the same pytest command. Use it when the laptop lacks the right interpreter. Do not treat a remote green run as proof about production data.

Review the diff before you trust the remote result. Reject any edits that land outside the order_status.py file. Reject any test change, alias deletion, or second helper.

If the draft is wider than the allowed cut, discard it. Ask again with the same constraints, or write the helper yourself. The table, not the draft, decides whether to merge.

Step 9. Accept the patch only after three checks

Count new functions, and allow only fold_status in this patch. Count changed files, and allow only the production module. Compare the alias literals with the parent revision.

A check fails if review disappears or payed changes meaning. A check fails if the log format drops the original text. Stop and revert the patch when any check fails.

Map the fixture onto a messy module

Start from one public function whose return value callers already depend on. Copy only synthetic inputs, or inputs your policy already allows in tests. Do not paste production payloads into the golden table.

Keep the file path stable while you extract the helper. A rename plus an extract hides the behavior diff. Land the rename only after this gate has passed.

Limits of the freeze

Characterization rows preserve defects until you change them on purpose. The payed alias and the review branch are current defects in this story. Shipping them longer is a product choice, not a testing victory.

A nine-row table can still miss rare inputs. Numbers other than twelve, bytes, and custom objects stay untested here. Expand the table only by executing new rows, then freezing them.

The free server may differ in Python version, locale, or installed plugins. Pin the test command in the patch notes so reviewers can repeat it. Do not upload secrets, customer orders, or private repository tokens.

A model can pass tests while making the code harder to read. Reject clever rewrites that hide the alias table. The goal is a smaller, reviewable diff, not a shorter file at any cost.

Who should skip this workflow

Skip it when the ticket requires a new label in the same change. Write the new expected rows first, then change the mapping. Calling that mix a pure extract will mislead reviewers.

Skip it when the entry point cannot run without production credentials. A borrowed database row is not a characterization fixture. Build a local input list, or do not claim the freeze is real.

Skip it for new code that has no callers and no current outputs. Ordinary unit tests should state the intended contract directly. There is nothing old to preserve in that case.

Skip it if your review policy forbids locking known defects even briefly. Some teams require the typo fix in the same pull request. Follow that policy instead of forcing this gate.

Commit order

Make the first commit contain only the golden test and the probe. Production behavior should be unchanged in that commit. Reviewers can run pytest before they read any move.

Make the second commit contain only fold_status and its call. The expected table must stay byte-for-byte stable across commits. Mention the baseline count in the commit message.

Leave alias cleanup, review-label redesign, and logger replacement for later commits. Each of those can change a pinned row. Each one needs its own test update before the code change.

After the gate

If you already use MonkeyCode, point the free model at this frozen suite. Keep the merge decision on the same nine passing tests.

Top comments (0)