DEV Community

Dakota Huang
Dakota Huang

Posted on

Object Identity Belongs in the Flag Overlay Contract

Split a flag overlay only after mutation rows are green. The caller dict is an input and an output. Unknown keys and boolean text belong in that same contract.

Move code only after those checks hold. This guide uses one small Python teaching module. The sample is a proposal, not a recorded production incident.

Run it locally before you trust any generated case. Treat every command below as a local check. Archive outputs before you edit the helper.

Why this helper breaks on extract

The helper named apply_flags mixes three jobs today. It reads defaults, parses text flags, and writes back into the same dict. A later extract that returns a new dict will change callers.

Those callers may still expect in-place edits. Boolean text is the second trap in this helper. A non-empty string is truthy in ordinary Python code.

A naive bool cast will keep that feature turned on. Empty strings and missing keys are different rows. They must not collapse into one shared case.

Rows that must stay green

Lock these rows before any extract work begins. Each row is a characterization check, not a wish. Michael Feathers named this style characterization testing in Working Effectively with Legacy Code.

Row Starting defaults Flag token Required result
1 debug false, limit 5 none same object, values unchanged
2 debug true debug=false boolean False, not the text false
3 debug true debug=FALSE boolean False after case fold
4 limit 5 limit= empty string stored
5 no region key region=us unknown key absent
6 limit 5 limit=10 string 10, not an int

Do not add numeric coercion during this pass. Numeric coercion would be a separate second change. The current contract keeps numeric text as strings.

1. Place the messy helper in one module

Create flag_overlay.py with the behavior you intend to lock. The function below is an unexecuted teaching sample. It is not shipped as a library release.

def apply_flags(defaults, raw_flags):
    """Mutate defaults from key=value tokens and ignore unknown keys."""
    parsed = {}
    for token in raw_flags:
        if "=" not in token:
            continue
        key, value = token.split("=", 1)
        parsed[key] = value
    for key, value in parsed.items():
        if key not in defaults:
            continue
        if isinstance(defaults[key], bool):
            defaults[key] = value.strip().lower() in {"1", "true", "yes", "on"}
        else:
            defaults[key] = value
    return defaults
Enter fullscreen mode Exit fullscreen mode

Notice the return value is the same object. Callers can ignore the return and still see edits. A pure extract must preserve that alias unless tests allow a break.

2. Lock identity and boolean text

Put the tests in test_flag_overlay.py beside the helper. Use unittest from the Python standard library only. The standard library page for unittest is the runner reference.

No third-party plugin is required for this pass. Read the Python unittest documentation before you change this runner. Keep the test file free of network calls.

import unittest
from flag_overlay import apply_flags

class OverlayContractTest(unittest.TestCase):
    def test_same_object_when_no_flags(self):
        defaults = {"debug": False, "limit": "5"}
        result = apply_flags(defaults, [])
        self.assertIs(result, defaults)
        self.assertEqual(defaults["debug"], False)
        self.assertEqual(defaults["limit"], "5")

    def test_false_text_becomes_boolean_false(self):
        defaults = {"debug": True, "limit": "5"}
        apply_flags(defaults, ["debug=false"])
        self.assertIs(defaults["debug"], False)

    def test_upper_false_text_is_also_false(self):
        defaults = {"debug": True, "limit": "5"}
        apply_flags(defaults, ["debug=FALSE"])
        self.assertIs(defaults["debug"], False)

    def test_empty_value_is_stored_for_text_fields(self):
        defaults = {"debug": False, "limit": "5"}
        apply_flags(defaults, ["limit="])
        self.assertEqual(defaults["limit"], "")

    def test_unknown_key_is_ignored(self):
        defaults = {"debug": False, "limit": "5"}
        apply_flags(defaults, ["region=us"])
        self.assertNotIn("region", defaults)

    def test_numeric_text_stays_a_string(self):
        defaults = {"debug": False, "limit": "5"}
        apply_flags(defaults, ["limit=10"])
        self.assertEqual(defaults["limit"], "10")
        self.assertIsInstance(defaults["limit"], str)
Enter fullscreen mode Exit fullscreen mode

The assertIs check on False targets the singleton, not equality. That check stops an accidental string False result. Keep one assertion theme inside each test method.

Mixed themes can hide the first real break. A failed boolean row should name only that row. Repair the fixture before you touch production lines.

3. Capture the baseline before any move

Run the module with python or python3 from the repo root. Capture both stdout and the process exit code. Do not format the log before you archive it.

python -m unittest test_flag_overlay.OverlayContractTest > overlay_baseline.txt 2>&1
echo $? > overlay_exit.txt
Enter fullscreen mode Exit fullscreen mode

A clean run prints OK and stores exit code zero. If a row fails, stop the extract immediately. Fix the fixture or record current behavior in the test.

Do not correct production code in the same commit as the move. The command above is a workflow sketch only. The exit capture uses bash or zsh syntax, not cmd.

Compare file bytes, not a pasted chat transcript. Store both files next to the test module. Those files are the contract snapshot for this pass.

4. Extract parsing and leave mutation alone

Extract parsing into a function named parse_tokens. Leave the mutation inside apply_flags for this pass. That split counts as one change by itself.

Returning a fresh copy would be a second change. Do not take both edits on the same day. Call sites should not need a rewrite yet.

def parse_tokens(raw_flags):
    parsed = {}
    for token in raw_flags:
        if "=" not in token:
            continue
        key, value = token.split("=", 1)
        parsed[key] = value
    return parsed

def apply_flags(defaults, raw_flags):
    parsed = parse_tokens(raw_flags)
    for key, value in parsed.items():
        if key not in defaults:
            continue
        if isinstance(defaults[key], bool):
            lowered = value.strip().lower()
            defaults[key] = lowered in {"1", "true", "yes", "on"}
        else:
            defaults[key] = value
    return defaults
Enter fullscreen mode Exit fullscreen mode

Callers of apply_flags stay untouched by this cut. New callers can use parse_tokens without inheriting mutation. That separation is the point of the smallest cut.

5. Diff the class output, not a memory of it

Repeat the same command after the extract lands. Diff the two captured text files after that. A green extract has no assertion text delta.

python -m unittest test_flag_overlay.OverlayContractTest > overlay_after.txt 2>&1
echo $? > overlay_exit_after.txt
diff -u overlay_baseline.txt overlay_after.txt
diff -u overlay_exit.txt overlay_exit_after.txt
Enter fullscreen mode Exit fullscreen mode

Expect both diffs to be empty on a safe move. A changed test count means the runner discovered a new file. Isolate the command to this module if that happens.

Test name order can also change the captured text. Pin the command to one test class if your tree is noisy. Use the narrower command shown above for both runs.

Then the diff compares that one class only. Keep the class path stable across both captures. A renamed class will look like a behavior change.

Decide the next edit from the diff

Use this table before you open a second commit. Give each observed signal exactly one next decision. Do not bundle a policy change into the extract commit.

Signal after the diff Allowed next edit Hold back
Empty diff and exit zero Keep the parse extract Number coercion
Boolean row fails Restore the old branch Key renames
Unknown key appears Put the ignore check back New warning logs
Object id changed Return the same dict Caller rewrites
Whitespace-only diff Normalize the runner Any production edit

Whitespace-only diffs are a test harness issue. They are not proof that the overlay is safer. Fix the runner command, then re-check the files.

Draft extra rows, then rerun them yourself

MonkeyCode free model access can draft extra boundary rows. Disclosure: This article was prepared as part of MonkeyCode's product outreach. A free server option can rerun the same unittest command.

Ask the model for extra cases, not a full rewrite. A useful prompt lists the six locked rows and forbids new coercion. Paste the helper and require tests to call apply_flags only.

Reject any draft that changes production code in the same patch. Review every drafted assertion by hand before merging. A model can invent a key the function already ignores.

A model can also treat the text false as a string. That mistake is why the boolean rows exist here. Keep the draft in a scratch file until the local suite agrees.

If you use the free server, run the pinned class command there. Copy the after file back to the laptop tree. Diff that file against the laptop baseline text.

A different Python minor version can change failure text. Pin the same minor version before you call the files equal. Do not treat a warning line as an overlay regression.

Limits you should state in the review

This pass does not prove thread safety at all. Two callers sharing one defaults dict can still race. The tests run in one process and one thread.

They will stay green even while a race exists. Shared defaults need a lock the tests do not provide. Do not cite this suite as a concurrency proof.

Quoted digit tokens currently store the quote marks. Tokens without an equals sign are skipped silently here. Add a row for either edge before you change it.

Duplicate tokens keep the last write in this sample. The sample does not record the first write. Add that row before you change merge order.

Silent ignore of unknown keys may be wrong for a strict CLI. Changing that policy is a product edit, not an extract. Write a failing row before you start rejecting unknown keys.

These teaching files were not executed in a published CI log. Run them on your machine before you cite a green result. An exit code of zero on your machine is the result that counts.

Who should leave this cut alone

Skip this cut if callers need a fresh dict every time. Identity tests would then freeze the wrong contract. Write the copy behavior first, with its own rows.

Skip the hosted path if flags can carry secrets. Do not send raw tokens to a hosted model or a shared server. Use local fixtures with fake values instead of live tokens.

Skip the extract if the helper is already pure and covered. An extra extract adds names without reducing risk. Spend the change budget on a failing production row.

Skip the diff ritual if you cannot pin the Python version. Diff noise will then look like a real regression. Fix the runner before you move any lines.

What to keep after the cut

Lock mutation, boolean text, empty values, and unknown keys first. Extract parse_tokens only after the class diff is empty. Leave number coercion and copy semantics for a later commit.

Each of those later edits is a separate contract change. If a free server session is already open, rerun the pinned class there. Diff the captured text against your laptop file.

Keep the green characterization file in the repo either way. That file is the evidence for the next reviewer. Do not replace that file with a chat summary.

Top comments (0)