DEV Community

Dakota Huang
Dakota Huang

Posted on

One Coercion Table Moves Only After Twelve Export Rows Hold

One Coercion Table Moves Only After Twelve Export Rows Hold

One coercion table moves only after twelve export rows hold. A messy mapper stays untouched until those frozen rows match. The smallest safe change is one pure coercion table.

The problem in one pass

Export helpers grow by one silent exception at a time. A null note becomes an empty string in one branch. A numeric string becomes an integer minor unit in another.

Callers then depend on both accidents without a written contract. A broad rewrite looks faster than a pinned fixture run. It also hides which accident a downstream report still needs.

Pin the current map before you touch the helper body. Move one coercion table only after those rows stay green. Treat any extra cleanup as a separate, later change.

What this sample freezes

The sample below is a proposed export mapper only. It is not a production trace from a named system. Expected rows are fixtures, not measured customer data.

Twelve input rows cover branches that usually drift during cleanup. Each row pairs one input dict with one expected dict. The test fails if key order, types, or null policy change.

Paste the helper and the test into a scratch project first. Run them before you adapt names to your repository. If your real helper differs, record its output, then edit fixtures.

Step 1. List the branches before you edit

Write the branch list before you open the helper for edits. Do not start this pass with a cleanup or a rename. The list below is the contract for this one change.

  1. A missing key stays absent, and it does not become null.
  2. A null note becomes an empty string in the output row.
  3. Boolean true becomes the text 1, and false becomes 0.
  4. An integer cent value is copied through as its decimal text.
  5. A numeric string of cents is stripped, then parsed as int.
  6. Unknown extra keys are dropped and never copied through.
  7. A whitespace-only note becomes an empty output string.
  8. A negative integer cent value stays negative in the output.
  9. A list of tags becomes one pipe-joined output string.
  10. A nested note dict becomes the token ERR_NESTED, not fields.
  11. A non-numeric cent value becomes the token ERR_CENTS.
  12. Header order follows a fixed column list, not insertion order.

If a branch is absent from the fixtures, leave that branch alone. Do not fix an unlisted branch inside this same change. Double-minus cent strings are one such unlisted gap here.

Step 2. Freeze outputs with one table test

Keep the mapper impure while you record the current outputs. Call it the same way production code calls it today. Assert the whole row, not a single favorite field.

The helper below is a proposal sketch you can run. It is not a decompiled capture from a live service. Boolean checks come before integer checks because bool subclasses int.

COLUMNS = ("sku", "qty", "cents", "flag", "note", "tags")

def map_row(raw):
    out = {}
    if "sku" in raw and raw["sku"] is not None:
        out["sku"] = str(raw["sku"]).strip()
    if "qty" in raw and raw["qty"] is not None:
        out["qty"] = str(int(raw["qty"]))
    if "cents" in raw and raw["cents"] is not None:
        value = raw["cents"]
        if isinstance(value, bool):
            out["cents"] = "ERR_BOOL"
        elif isinstance(value, int):
            out["cents"] = str(value)
        elif isinstance(value, str) and value.strip().lstrip("-").isdigit():
            out["cents"] = str(int(value.strip()))
        else:
            out["cents"] = "ERR_CENTS"
    if "flag" in raw and raw["flag"] is not None:
        if raw["flag"] is True:
            out["flag"] = "1"
        elif raw["flag"] is False:
            out["flag"] = "0"
        else:
            out["flag"] = "ERR_FLAG"
    if "note" in raw:
        note = raw["note"]
        if note is None or (isinstance(note, str) and note.strip() == ""):
            out["note"] = ""
        elif isinstance(note, dict):
            out["note"] = "ERR_NESTED"
        elif isinstance(note, str):
            out["note"] = note.strip()
        else:
            out["note"] = "ERR_NOTE"
    if "tags" in raw and raw["tags"] is not None:
        tags = raw["tags"]
        if isinstance(tags, list):
            out["tags"] = "|".join(str(item) for item in tags)
        else:
            out["tags"] = "ERR_TAGS"
    ordered = {}
    for key in COLUMNS:
        if key in out:
            ordered[key] = out[key]
    return ordered
Enter fullscreen mode Exit fullscreen mode

Save the helper as export_map.py in the repository root. Run unittest from that root so the import resolves. The sample assumes no package prefix on the module name.

import unittest
from export_map import map_row

CASES = [
    {
        "name": "missing-key-stays-absent",
        "raw": {"sku": "A1", "qty": 2},
        "out": {"sku": "A1", "qty": "2"},
    },
    {
        "name": "null-note-to-empty",
        "raw": {"sku": "A1", "note": None},
        "out": {"sku": "A1", "note": ""},
    },
    {
        "name": "bool-flag-true",
        "raw": {"sku": "A1", "flag": True},
        "out": {"sku": "A1", "flag": "1"},
    },
    {
        "name": "bool-flag-false",
        "raw": {"sku": "A1", "flag": False},
        "out": {"sku": "A1", "flag": "0"},
    },
    {
        "name": "int-cents-unchanged",
        "raw": {"sku": "A1", "cents": 250},
        "out": {"sku": "A1", "cents": "250"},
    },
    {
        "name": "string-cents-stripped",
        "raw": {"sku": "A1", "cents": " 250 "},
        "out": {"sku": "A1", "cents": "250"},
    },
    {
        "name": "unknown-key-dropped",
        "raw": {"sku": "A1", "debug": "x"},
        "out": {"sku": "A1"},
    },
    {
        "name": "blank-note-to-empty",
        "raw": {"sku": "A1", "note": "  \t"},
        "out": {"sku": "A1", "note": ""},
    },
    {
        "name": "negative-cents-stay-negative",
        "raw": {"sku": "A1", "cents": -40},
        "out": {"sku": "A1", "cents": "-40"},
    },
    {
        "name": "tags-join-with-pipe",
        "raw": {"sku": "A1", "tags": ["red", "blue"]},
        "out": {"sku": "A1", "tags": "red|blue"},
    },
    {
        "name": "nested-note-error-token",
        "raw": {"sku": "A1", "note": {"a": 1}},
        "out": {"sku": "A1", "note": "ERR_NESTED"},
    },
    {
        "name": "bad-cents-token-and-order",
        "raw": {"note": "ok", "cents": "1.5", "sku": "A1"},
        "out": {"sku": "A1", "cents": "ERR_CENTS", "note": "ok"},
    },
]

class ExportMapFreezeTest(unittest.TestCase):
    def test_twelve_rows_hold(self):
        for case in CASES:
            got = map_row(case["raw"])
            self.assertEqual(case["out"], got, case["name"])
            self.assertEqual(list(case["out"]), list(got), case["name"])
Enter fullscreen mode Exit fullscreen mode

Case twelve checks a bad cent token and key order together. Dict equality ignores key order, so the test also compares key lists. The expected key list is sku, then cents, then note.

Python 3.7 and later preserve insertion order for regular dicts. The order assertion relies on that language rule, not on a sort. If you run an older interpreter, replace dicts with an ordered list of pairs.

Run one command, and do not add fixtures after a red row. A red row is evidence, so save the actual dict beside its name. Update a fixture only after a human confirms today's helper.

python -m unittest tests/test_export_map_freeze.py -v
Enter fullscreen mode Exit fullscreen mode

A model draft is not that confirmation, even when the draft looks neat. Compare the draft to a local run of the current helper. Keep the local dict when the two outputs disagree.

This sample is unexecuted in this draft, so run it before you trust it. The expected dicts were derived by inspection of the helper above. If a local run disagrees, keep the local dict and fix the writeup.

Step 3. Decide the one allowed move

Use a small gate before you allow any production edit. If a row is red for an unknown reason, stop the change. If all twelve rows match, one extract is the only allowed move.

Gate Evidence required Allowed next edit Forbidden in the same diff
Red, unexplained Actual row saved next to the name None Any helper edit
Red, fixture typo Human diff of one expected cell That fixture cell only Mapper edit
Green, twelve rows Full unittest log from the local tree Extract one coercion table Behavior change, new column, rename
Green, then flaky Two identical local reruns None until the pair is stable Timing edits and sleeps

The allowed edit is mechanical and stays inside one new table. Move type rules into that table, and leave call order alone. Leave error tokens unchanged until a later, separate review.

def coerce_cents(value):
    if isinstance(value, bool):
        return "ERR_BOOL"
    if isinstance(value, int):
        return str(value)
    if isinstance(value, str) and value.strip().lstrip("-").isdigit():
        return str(int(value.strip()))
    return "ERR_CENTS"

COERCE = {
    "sku": lambda v: str(v).strip(),
    "qty": lambda v: str(int(v)),
    "cents": coerce_cents,
    "flag": lambda v: "1" if v is True else "0" if v is False else "ERR_FLAG",
    "note": lambda v: (
        "" if v is None or (isinstance(v, str) and v.strip() == "")
        else "ERR_NESTED" if isinstance(v, dict)
        else v.strip() if isinstance(v, str)
        else "ERR_NOTE"
    ),
    "tags": lambda v: (
        "|".join(str(item) for item in v) if isinstance(v, list) else "ERR_TAGS"
    ),
}
Enter fullscreen mode Exit fullscreen mode

Wire map_row through COERCE without changing the presence checks. A missing key must still stay absent after the extract. A present null note must still become an empty string.

Do not fix the string cent parser in this diff. The lstrip check accepts some odd forms and can raise on others. That gap stays unlisted until a new named case exists.

Step 4. Use a free model only as a drafter

Disclosure: This article was prepared as part of MonkeyCode's product outreach. Operator notes supplied for this draft say MonkeyCode offers free model access. The same notes say a free server option is available.

Those notes do not state quotas, model names, hardware, or duration. This workflow uses free model access only to draft candidate rows. This workflow uses the free server only to run a copied tree.

Neither choice is a source of expected values. Give the model the branch list and the current function text. Ask for failing test names, not for a rewritten module.

Reject any draft that changes an expected value without current output. A practical split keeps merge authority on the local tree. The local editor holds the real helper and the approved fixtures.

The free server holds a copy of the helper plus the draft. Copy back only a test file after you inspect the diff. Re-run that file locally before you trust any server result.

Do not copy back a rewritten mapper from the server copy. The second local run is the gate for this workflow. A green run on a free server is not the gate.

Server copies drift, and your tree is the one callers use. Label server logs as untrusted notes, not as fixtures. Paste a server dict into the test only after a local match.

If the copy cannot import the helper, fix the copy, not the rules. Treat an import error as a copy problem, not as a product verdict. Do not weaken a fixture to make a broken copy pass.

python -m unittest tests/test_export_map_freeze.py -v
diff -u /tmp/server-draft/test_export_map_freeze.py tests/test_export_map_freeze.py
python -m unittest tests/test_export_map_freeze.py -v
Enter fullscreen mode Exit fullscreen mode

Step 5. Review the diff with a line budget

Count the production diff after the coercion extract lands. A safe move stays inside one new function and one dict. Imports should not grow, and public names should not change.

git diff --stat -- export_map.py
git diff -U3 -- export_map.py
Enter fullscreen mode Exit fullscreen mode

Stop if the stat output shows a second production file. Stop if a string token changed inside the helper body. Stop if a test file changed in the same commit as the extract.

Split those edits into separate commits before you continue. Rerun the twelve-row file after each split commit lands. A green log from before the split does not cover the new tip.

Track three numbers in the review note, not a prose impression. Count production files touched, assertion lines edited, and tokens changed. The target is one file and zero assertion edits.

The same target also includes zero error-token edits. Write those three counts in the commit message body. Reviewers should reject the diff when any count misses.

What the twelve rows do not prove

These rows do not prove the mapper is correct for finance. They prove today's outputs are stable for those twelve inputs. A locked bug remains a bug after the freeze goes green.

The string cent parser is not a decimal policy. It does not define rounding, currency scale, or overflow. Do not cite these fixtures as evidence for a money rule.

These rows do not cover concurrency, encoding, or CSV quoting. Add those risks only as new named cases in a later change. Do not fold them into the coercion extract.

These rows do not measure model quality or server speed. No latency number, token count, or pass rate is claimed here. A free run that drafts ten bad rows is still useful evidence.

Python treats bool as a subclass of int, so order matters. That is why the cent branch checks bool before int. Removing that order would map True to the text 1.

Who should skip this workflow

Skip this flow when the helper writes payments or deletes rows. Skip it when the helper calls a live vendor during mapping. Characterization still matters there, but not against live side effects.

Skip this flow when no human can bless the actual output. A free model cannot supply that blessing for you. Skip this flow when the goal is a new column or rounding law.

Write a behavior test with a new expected row for that goal. Do not disguise a policy change as a small refactor. Skip this flow when fixtures include customer payloads or secrets.

Signed URLs and tokens do not belong in characterization files. Use synthetic rows like the twelve cases shown above. Skip the server copy if your policy forbids leaving the laptop.

One next check

If the twelve rows are green, extract the table and rerun the file. If any case name flips red, revert the extract immediately. Leave the fixtures in place, because that revert is the point.

Readers can run the draft copy on MonkeyCode's free server option. Keep the merge decision on the machine that owns the real tree.

Top comments (0)