Lock the reject ledger before you extract a parser. A green ledger proves the split did not change triage. The smallest safe change is one pure classifier, not a rewrite.
Why this extract is the whole job
Messy ingest functions hide three jobs in one loop. They skip blanks, parse JSON, and record rejects. They also drop duplicate ids while building a kept list.
Splitting all of that at once hides regressions. This note covers one Python module and one test file. The target is a line ingest helper in a messy repo.
No network call belongs inside this characterization change. The helper takes a list of strings and returns two lists. You can replay that pair without a service.
The contract you must freeze
Read the current loop before you edit it. Write the observed outcomes into a frozen table. Treat that table as the oracle for this change.
| case | input shape | kept ids | rejects |
|---|---|---|---|
| blanks | empty and spaces | none | none |
| one row | object id a | a | none |
| bad json | lone brace | none | (1, bad_json) |
| missing id | object without id | none | (1, missing_id) |
| non-object | JSON list | none | (1, missing_id) |
| duplicate | two objects id a | a | (2, dup_id) |
| blank then bad | blank, then brace | none | (2, bad_json) |
| null id key | object id null | null | none |
That last row is easy to fix by accident. The key exists, so the current loop keeps it. Do not change that rule in this extract.
Line numbers count every input row, including blanks. A blank is a skip, not a reject. A later bad line keeps its original index.
The snippets below are unexecuted examples for this workflow. Run them in your repo before you trust the pass count. Treat a local green run as the only acceptance signal.
Step 1. Pin the messy function
Save the current behavior in a module named ingest_legacy.py. Do not clean it while you copy it. The characterization tests import this module before any edit.
import json
def ingest(lines):
seen = set()
kept = []
rejects = []
for idx, raw in enumerate(lines, start=1):
text = raw.strip()
if text == "":
continue
try:
obj = json.loads(text)
except json.JSONDecodeError:
rejects.append((idx, "bad_json"))
continue
if not isinstance(obj, dict) or "id" not in obj:
rejects.append((idx, "missing_id"))
continue
if obj["id"] in seen:
rejects.append((idx, "dup_id"))
continue
seen.add(obj["id"])
kept.append(obj)
return kept, rejects
This function is deterministic for a fixed list. It does not read the clock, the cwd, or the environment. That is why this ledger can be a pure contract.
Step 2. Encode the table as tests
Put every case in tests/test_ingest_contract.py before editing production code. Each case asserts kept ids and the full reject list. Assert the reject tuple, not a substring of a message.
import pytest
from ingest_legacy import ingest
CASES = [
("blanks", ["", " "], [], []),
("one", ['{"id":"a","n":1}'], ["a"], []),
("bad", ["{"], [], [(1, "bad_json")]),
("noid", ['{"name":"a"}'], [], [(1, "missing_id")]),
("list", ["[1]"], [], [(1, "missing_id")]),
("dup", ['{"id":"a","n":1}', '{"id":"a","n":2}'], ["a"], [(2, "dup_id")]),
("blank_bad", ["", "{"], [], [(2, "bad_json")]),
("null_id", ['{"id": null}'], [None], []),
]
@pytest.mark.parametrize("name,lines,ids,rejects", CASES)
def test_ledger(name, lines, ids, rejects):
kept, got = ingest(lines)
assert [row.get("id") for row in kept] == ids
assert got == rejects
The duplicate case also locks first-seen payload fields. Add one assertion that the first kept n equals 1. The second row must not overwrite the first n.
def test_first_payload_wins():
lines = ['{"id":"a","n":1}', '{"id":"a","n":2}']
kept, rejects = ingest(lines)
assert kept[0]["n"] == 1
assert rejects == [(2, "dup_id")]
Run the full suite before any production code edit. Record the pass count in the commit message. A later run must show the same count.
python -m pytest tests/test_ingest_contract.py -q
If a case fails, the table was wrong, not the code. Fix the expected row and rerun the same file. Do not edit ingest until every listed case passes.
Step 3. Hash the ledger for a second check
A pass count can hide a swapped fixture. Hash the serialized outcomes for the eight cases. Store the digest in the test module as a comment.
Do that only after the first green run. Build pairs from the same CASES loop as id lists plus rejects. Print the digest once and copy the hex.
import hashlib
import json
def ledger_digest(pairs):
blob = json.dumps(pairs, sort_keys=True, separators=(",", ":"))
return hashlib.sha256(blob.encode("utf-8")).hexdigest()
# Fill EXPECTED_DIGEST after the first local green run.
# Leave it unset until that hex is printed locally.
EXPECTED_DIGEST = None
Paste that hex into the test file as EXPECTED_DIGEST. Assert digest equality on every later test run. Do not update the digest to match a new extract.
Update it only when the product owner accepts a behavior change. This extract is not a behavior acceptance change. A matching hex is the merge gate for this commit.
Step 4. Draft the classifier outside the oracle
Free model access can draft the pure function from the loop text. Disclosure: This article was prepared as part of MonkeyCode's product outreach. Use that access to propose classify_line and nothing else.
Reject any draft that also rewrites dedupe, logging, or file output. A free server option can run the same pytest command on a clean checkout. Compare that remote digest with your local hex.
A digest mismatch stops the extract before any merge. Do not fix the server result by editing the oracle. The free model draft is not the oracle.
The table you wrote in step 2 is the oracle. If the draft adds a null-id reject, discard the draft. That draft would change locked table row eight.
Free access here means model use and a server option only. It does not name a model, a quota, or an uptime promise. Do not plan a release around either availability claim.
Step 5. Apply the smallest edit
Move only the per-line decision into a pure function. Leave the seen set, kept list, and rejects inside ingest. Duplicate handling stays in the loop because it needs cross-line state.
def classify_line(text):
text = text.strip()
if text == "":
return ("skip", None)
try:
obj = json.loads(text)
except json.JSONDecodeError:
return ("bad_json", None)
if not isinstance(obj, dict) or "id" not in obj:
return ("missing_id", None)
return ("ok", obj)
def ingest(lines):
seen = set()
kept = []
rejects = []
for idx, raw in enumerate(lines, start=1):
kind, obj = classify_line(raw)
if kind == "skip":
continue
if kind != "ok":
rejects.append((idx, kind))
continue
if obj["id"] in seen:
rejects.append((idx, "dup_id"))
continue
seen.add(obj["id"])
kept.append(obj)
return kept, rejects
Re-run the same pytest command after the edit. The pass count must match the step 2 count. The ledger digest must match the step 3 hex.
python -m pytest tests/test_ingest_contract.py -q
git diff --stat -- ingest_legacy.py
Then run a diff stat and confirm one module changed. If the diff touches tests, you moved the oracle. Revert those test edits before you commit anything.
Tests should change only when the contract changes. This commit does not change the contract at all. Keep the test file byte-stable after the digest comment.
Decision matrix before you touch code
| signal | decision |
|---|---|
| suite fails on the untouched copy | correct the table |
| draft rejects a null id | discard the draft |
| remote digest differs | stop and compare fixtures |
| ticket demands a new dup rule | do not start |
| diff includes test expectation edits | revert those edits |
Use the matrix as a stop list, not as style advice. One matching row is enough to halt the extract. Two matching rows still mean halt, not debate.
Three failures this ledger catches
A swapped duplicate payload fails the n equals 1 check. The second JSON object must not replace the first. That catch is why the extra assertion exists.
A blank line counted as a reject fails the blank rows. Index drift then breaks the blank-then-bad case. Both of those checks must stay green together.
A classifier that reads seen is no longer a pure extract. The diff should show no seen reference inside classify_line. Cross-line state stays in the loop on purpose.
What this extract does not prove
A green ledger does not prove the helper is fast. It does not prove JSON numbers keep type beyond these rows. It does not cover files, sockets, or partial reads.
Those inputs sit fully outside this locked contract. json.loads accepts duplicate keys by keeping the last key. This table does not pin that parser rule.
Add a case before you rely on it. Do not sneak that case in after the extract. Very long lines and non-UTF8 bytes are out of scope.
These characterization tests pass Python str lines only. If your caller passes bytes, stop and write a new contract first. Build that contract in a separate change with its own digest.
Who should skip this workflow
Skip it when the ticket asks for a new dup policy. A characterization test would freeze the old policy. You would then fight your own green suite.
Skip it when the line source is not replayable. A live socket cannot feed the same eight cases twice. Capture a fixture first, then lock the ledger.
Skip it when sample lines contain secrets or customer payloads. Use synthetic ids such as a and n. A free server run is still a run of your fixture text.
Skip it when you cannot run pytest locally. A remote pass without a local digest is not a lock. You need both results before you merge.
Review list before you merge
Confirm eight table rows still match the assertions. Confirm first payload field n stays 1 on duplicate id a. Confirm blank lines do not appear in rejects.
Confirm bad JSON after a blank keeps index 2. Confirm classify_line does not read the seen set. Confirm the diff stat shows no test-oracle edits.
Confirm the local digest equals the free-server digest. Confirm no second extract landed in the same commit. If any item fails, revert the function edit.
Keep the characterization tests exactly as they stand. Fix the extract itself, not the locked ledger. Repeat the pytest command until the hex matches.
Where a second change would start
The next safe extract is the duplicate check, not a printer. Give it a new table for three ids and two dups. Run that table to green before you move code.
Stop the work after that function returns cleanly. Do not combine the printer, the file writer, and the classifier. Each of those jobs needs its own green ledger.
This commit should end at the classify_line function. If you have a free MonkeyCode server, run the pinned pytest command there. Compare the digest, then review the diff yourself before you publish the branch.
Top comments (0)