Picture a billing worker that skipped invoice generation after a weekend deploy because a blank region string replaced a valid environment value. In that scenario, the on-call engineer expected missing keys and blank values to behave the same way under overlay. They did not, and the same function treated a JSON null as an explicit delete of that key. File reads, environment scans, and CLI parsing lived in one module, so a broad cleanup looked useful and unsafe.
The notes below describe a characterization-first path for that class of settings merger, using a labeled composite module. Nothing in this draft reports a private employer incident, a measured outage length, or any customer count. The tests pin today's behavior before anyone moves code, including behavior that a later change may deliberately reject. Product assistance appears only after that contract exists, and the outreach disclosure sits beside the first mention.
What must stay stable during the extract
A safe extract preserves outputs for every input class you can name, even when those outputs look wrong. The merger currently walks sources in file, environment, then CLI order, and later writes replace earlier writes. A null value deletes the key, while an empty string, zero, and an empty list remain stored values. Moving I/O, renaming keys, and fixing null handling in the same diff would hide which edit changed behavior.
Decision table before any code move
Lock the contract in a table before you touch the module, because prose arguments drift during review. Each row names three source states and the exact merged result you observed, not the result you prefer. You should record absent keys, empty strings, JSON null, numeric zero, and empty collections as distinct states. If two reviewers disagree on a row, rerun the current function and paste the observed output into the table.
| Case | File value | Env value | CLI value | Observed merge |
|---|---|---|---|---|
| later value wins | region=us | region=eu | absent | region=eu |
| empty string wins | region=us | region=empty string | absent | region=empty string |
| null deletes | region=us | region=null | absent | region absent |
| zero is kept | retries=3 | retries=0 | absent | retries=0 |
| empty list kept | tags=[a] | tags=[] | absent | tags=[] |
| cli null deletes | region=us | region=eu | region=null | region absent |
| delete then set | region=null | region=eu | absent | region=eu |
Composite module under test
Label this module as an unexecuted teaching example until you paste it into a scratch repository and run it. It intentionally keeps the awkward rules so the tests describe the code you have, not the code you want. Do not import production secrets, live endpoints, or customer fixtures into this small characterization harness at all. The function accepts plain dictionaries so characterization does not depend on disk layout or the process environment.
Messy function to pin
# scratch/overlay.py
# Labeled composite. Run it before you treat any row as evidence.
def overlay_settings(file_cfg, env_cfg, cli_cfg):
merged = {}
for source in (file_cfg, env_cfg, cli_cfg):
for key, value in source.items():
if value is None:
merged.pop(key, None)
else:
merged[key] = value
return merged
Characterization tests that refuse silent drift
These tests call the current function and compare full dictionaries, because a partial assert can miss deleted keys. A missing key and a stored null are different outcomes, so the expected object must not contain a placeholder. Keep one assertion per case name, and let that case name match the corresponding decision table row. If a test fails on the first run, fix the expectation to match observed output before you edit production code.
# tests/test_overlay_characterize.py
import pytest
from overlay import overlay_settings
CASES = [
("later_value_wins", {"region": "us"}, {"region": "eu"}, {}, {"region": "eu"}),
("empty_string_wins", {"region": "us"}, {"region": ""}, {}, {"region": ""}),
("null_deletes", {"region": "us"}, {"region": None}, {}, {}),
("zero_is_kept", {"retries": 3}, {"retries": 0}, {}, {"retries": 0}),
("empty_list_kept", {"tags": ["a"]}, {"tags": []}, {}, {"tags": []}),
(
"cli_null_deletes",
{"region": "us"},
{"region": "eu"},
{"region": None},
{},
),
("delete_then_set", {"region": None}, {"region": "eu"}, {}, {"region": "eu"}),
]
@pytest.mark.parametrize(
"name,file_cfg,env_cfg,cli_cfg,expected",
CASES,
ids=[row[0] for row in CASES],
)
def test_overlay_characterizes_current_contract(
name, file_cfg, env_cfg, cli_cfg, expected
):
del name # parametrize ids already label the failure
assert overlay_settings(file_cfg, env_cfg, cli_cfg) == expected
Commands for a local proof
Run the suite from a clean shell so inherited variables do not masquerade as deliberate environment input. The sample function does not read the process environment, but the next extract might, and the habit should start now. Save the command output with the commit that introduces the tests, because a later green run is not the original evidence. If your laptop shell exports REGION or RETRIES, the clean-server rerun described later is the stricter check.
python -m venv .venv
. .venv/bin/activate
python -m pip install pytest
env -u REGION -u RETRIES python -m pytest -q tests/test_overlay_characterize.py
python -c 'from overlay import overlay_settings; print(overlay_settings({"region": "us"}, {"region": None}, {}))'
python -c 'from overlay import overlay_settings; print(overlay_settings({"region": "us"}, {"region": ""}, {}))'
git diff --stat -- overlay.py tests/test_overlay_characterize.py
git diff -- overlay.py
Smallest safe change after the pin
The smallest safe change extracts the loop body into a pure helper and leaves file, environment, and CLI loading untouched. Callers still pass three dictionaries in the same order, and null still deletes while empty strings still overwrite. That boundary is one function move plus an import update, which a reviewer can check against the unchanged test file. A second change may later redefine null, but only after the table marks those rows as intended breaks.
def apply_source(merged, source):
for key, value in source.items():
if value is None:
merged.pop(key, None)
else:
merged[key] = value
return merged
def overlay_settings(file_cfg, env_cfg, cli_cfg):
merged = {}
for source in (file_cfg, env_cfg, cli_cfg):
apply_source(merged, source)
return merged
What not to smuggle into the move
Adding deepcopy, sorting keys, or coercing empty strings to null would change observable results for some rows. Shared mutable values remain part of the current contract until a test proves callers rely on isolation. Sorting keys would change JSON dump order if a later caller serializes the dict without an explicit sort. Leave those ideas in a follow-up list, and keep their code out of this characterization-backed diff.
Review checklist before merge
- Confirm every decision-table row has a passing test whose expected value came from a run, not from a design discussion.
- Confirm the diff does not add file reads, environment reads, or argument parsing inside the new helper.
- Confirm falsy values other than null, including zero and empty collections, still survive a later source write.
- Confirm a planned behavior change is absent from this commit, or is listed as an explicit failing-test follow-up.
Where free model access and a free server fit
Once the table and tests exist, MonkeyCode's free model access can propose extra rows that your first pass missed. Disclosure: This article was prepared as part of MonkeyCode's product outreach. The operator supplied two availability claims for this draft: free model access, and a free server option for running work away from a laptop. This article does not name models, token quotas, hardware sizes, time limits, or benchmark scores, because those details were not verified here.
How to use the assistant without letting it rewrite the contract
Ask for candidate rows only, and require each proposal to cite a source cell and an expected dictionary. Reject any suggestion that corrects null into a stored value, unless you opened a separate behavior-change task. Paste accepted rows into the test file yourself, then run the suite, because an unrun suggestion is not evidence. Treat the open-source project as a helper for drafting and execution, not as a source of production truth.
Why a free server rerun matters for this bug
Environment leaks are part of the original failure mode, so a second run on a machine you did not customize is useful. A free server option lets you repeat the same pytest command outside a shell that already exports billing variables. Record the image or setup notes you can actually verify, and do not assume the remote box matches production libc, locale, or timezone.
If the remote result differs from the laptop result, stop the extract and reconcile the inputs before you trust either green log. If the characterization table is already green locally, one remote rerun is enough of a second opinion before you merge the helper. Keep that remote log next to the test file so a reviewer can see which command produced the green result.
Limitations and who should not use this path
This path assumes you can execute the current function on synthetic dictionaries and read stable, comparable outputs. Skip it when the merger calls a live network, a paid vendor, or a database you cannot stub without changing behavior. Also skip it when product policy already demands a null-semantics change in the same release, because characterization-then-extract would delay a required break. Free model access can invent plausible rows that never occur in your payloads, so every accepted row still needs a local run.
A free server option is a cleaner shell, not a production replica, and unverified hardware claims would only add false confidence. Do not use the remote run as evidence if you cannot record which command you executed and which test file you copied. Teams without permission to upload even synthetic fixtures should keep the suite on an approved internal runner instead. If you cannot explain absent, empty, and null to a reviewer in one table, you are not ready to extract the helper.
What this workflow deliberately leaves undone
The extract does not choose a new precedence policy, and it does not migrate callers to a typed settings object. It also does not prove thread safety, file encoding, or comment-preserving YAML edits, because those risks sat outside the pinned contract. A later change can replace null-delete with store-null, but it should flip specific table rows and show the red tests first. Until that change exists, the honest description of the system is the table you ran, not the overlay you wish you had.
Top comments (0)