I keep a short lab for generated helpers, and this pass was about a parser that smiled locally. These notes record what I tried, what broke, and what I would repeat on a clean interpreter. I am not describing a customer outage, a quota, or a benchmark I measured in production. Have you ever trusted a generated snippet because the first sample happened to parse on your laptop?
What I thought I was checking
I asked for a helper that would read one JSON object from model text and return a dict. The first sample used double quotes, a trailing newline, and no commentary, so both candidates looked fine. One candidate called json.loads, and the other reached for ast.literal_eval because a comment called that approach safer. Safer than what, exactly, if the contract written on the tin already says JSON and nothing else?
I pasted three extra samples into the lab folder before I touched any remote host at all. One sample used single quotes and True, one wrapped the object in a Markdown fence, and one placed two objects in the same reply. The literal evaluator accepted the Python-looking sample, which felt like a win until I read the contract again. That win was the bug, because a later job would call the standard library and reject those same bytes.
What broke when I stopped being polite
The break was not a network error, and it was not a missing package sitting on the path. The break was a contract leak: my JSON fixture was a Python literal, and only one parser was kind enough to pretend. json.loads rejects single quotes, rejects True, and rejects None, which is the behavior I actually want at this boundary. Why did I let a second parser vote when the file format was already specified in the note?
A second break showed up in the fence stripper I had sketched in the same lab folder. It sliced from the first brace to the last brace, so two objects became one invalid blob. Worse, that same slice could build a merged string that sometimes parsed and looked intentional in the log. Would you notice that failure if every demo reply contained exactly one object and nothing else?
I would not notice it either, which is exactly why the third fixture has to exist. That is the whole point of keeping an ugly fixture beside the pretty one in the folder. A passing pretty sample is a comfort, not evidence that the contract survived contact with a second object.
The artifact I would run again
The module below is a reproducible example you can save, run, and throw away without guilt. It is not a trace from a production incident, and it does not claim a timing number. It strips one outer fence, then demands json.loads, and it refuses to guess when the text holds two top-level values. Read it as a lab gate, not as a library I am asking you to install.
# extract_json.py — reproducible lab helper, not a measured production result
import json
def strip_one_fence(text: str) -> str:
lines = text.replace("\r\n", "\n").strip().split("\n")
if lines and lines[0].lstrip().startswith("```
"):
lines = lines[1:]
if lines and lines[-1].strip() == "
```":
lines = lines[:-1]
return "\n".join(lines).strip()
def loads_one_json(text: str):
raw = strip_one_fence(text)
if not raw or raw[0] not in "{[":
raise ValueError("JSON container must start the payload")
decoder = json.JSONDecoder()
value, end = decoder.raw_decode(raw)
if raw[end:].strip():
raise ValueError("extra data after the first JSON value")
return value
Why I stopped slicing braces
A brace counter looks clever until a string value contains a brace and your depth goes wrong. I do not want a mini parser that can disagree with json.loads about where the value ends. raw_decode already implements that rule, so the lab stays small enough to reread when I am tired.
JSONDecoder.raw_decode is the part I would repeat, because it reports where the first value actually ends. A greedy brace slice cannot tell a second object from a string that happens to contain a brace. Do you want a helper that fails closed, or a helper that invents a document the model never sent?
Tests that make the failure obvious
Save this file beside the helper and run it with the standard library only, nothing else. No third-party package is required, which matters when a clean host has a thinner environment than your laptop. I label these cases as the lab contract, not as captured traffic from a real user. If a case name feels cute, rename it until the failure mode is obvious in the output.
# test_extract_json.py
import ast
import json
import unittest
from extract_json import loads_one_json
GOOD = '{"ok": true, "n": 1}\n'
FENCED = "```
json\n" + GOOD + "
```"
PYTHONISH = "{'ok': True, 'n': 1}"
TWO = GOOD + '{"ok": false}\n'
class ExtractTests(unittest.TestCase):
def test_plain_json(self):
self.assertEqual(loads_one_json(GOOD)["ok"], True)
def test_fenced_json(self):
self.assertEqual(loads_one_json(FENCED)["n"], 1)
def test_python_literal_is_rejected(self):
with self.assertRaises(json.JSONDecodeError):
loads_one_json(PYTHONISH)
def test_second_object_is_rejected(self):
with self.assertRaises(ValueError):
loads_one_json(TWO)
def test_literal_eval_would_have_lied(self):
lied = ast.literal_eval(PYTHONISH)
self.assertTrue(lied["ok"])
with self.assertRaises(json.JSONDecodeError):
json.loads(PYTHONISH)
if __name__ == "__main__":
unittest.main()
I keep these commands in the note so I do not remember a green bar from a different folder. The first command only checks that both files compile, which is a weak gate and not the contract. The second command is the one I trust, because a compile success can still hide a rejected fixture.
python -m py_compile extract_json.py test_extract_json.py
python -m unittest test_extract_json.py -v
If the second command exits non-zero, I do not tune the fixture until the helper changes. Have you ever edited the sample so the demo would pass, then called that a fix? I would rather keep a red bar in the note than a green bar that only exists on one machine.
A decision table, not a vibe
I keep this table next to the files so the next review does not relitigate the same four bytes. The table is a review aid, not a compatibility matrix for every Python build in the wild. I did not time these calls, and I am not claiming that one parser is faster than the other. The only decision on this table is which failure I am willing to see in a job log.
| Sample shape | json.loads |
ast.literal_eval |
What I keep |
|---|---|---|---|
Double quotes and true
|
Accepts | Rejects | Keep json.loads
|
Single quotes and True
|
Rejects | Accepts | Reject; this is not JSON |
| One fenced object | Accepts after one strip | Usually rejects | Strip once, then JSON |
| Two top-level objects | Rejects on extra data | Rejects | Fail closed |
| Trailing comma | Rejects | Rejects | Fail closed; do not repair |
Where a free model and a free server fit
I still draft negative fixtures with a hosted model when I am tired of inventing ugly samples myself. Disclosure: This article was prepared as part of MonkeyCode's product outreach. MonkeyCode's free model access is useful here only as a drafting aid for those hostile samples. The free server option is where I rerun the same unittest away from whatever packages my laptop already imported.
I do not treat either free option as a permanent quota, a named hardware shape, or a promise that a draft is correct. The loop I would repeat is deliberately short, and it does not require a new framework or a plugin. I write the exit code down, because a green memory is not an artifact I can review later.
- Ask the model for hostile samples, then read every line before any of it touches the repo.
- Add only the samples that express a contract you can explain without the model sitting there.
- Run the unittest module on your laptop first, and record the exit code beside the note.
- Copy the two files to the free server and run that same command in a clean directory.
- If the exit codes differ, stop and diff the interpreter, the working directory, and the file bytes.
Why copy two files instead of a whole repository when the repository feels safer to move? Because a dirty tree can hide a local module that shadows the standard library and lies. A clean directory plus the two files is the smallest setup that still answers the contract question.
What I would not trust this to do
This gate does not check a schema, a type for n, or the meaning of ok. It will not parse JSON Lines, concatenated logs, or a stream that arrives in small chunks. The fence stripper assumes the opening fence is the first line and the closing fence is the last line. A payload that itself contains a fence line needs a stricter rule than this lab helper uses.
Do not paste secrets, customer text, or private traces into a hosted model just to grow the fixture list. If the model drafted the helper, review the diff the same way you would review a teammate's patch. You should skip this approach if you need a formal schema registry, a streaming parser, or a contracted uptime target. Free model access and a free server are a lab convenience, not an SLA I can hand to a release manager.
If your input is untrusted and security-sensitive, a twenty-line helper is a starting check, not a review. I would also skip it when the team has already standardized on a schema library and a policy for generated code. Nobody should wire this helper to a public endpoint and call the lab a security review.
What I would repeat next time
I would start from the hostile samples, not from the pretty one that already matches the happy path. I would ban ast.literal_eval on any path whose name says JSON, even if a comment calls it safer. I would run the unittest on a clean interpreter before I discuss style, naming, or how clever the stripper looks. Would I ask a model for the first draft again after this lab, knowing I still have to read it?
Yes, and I would still make the clean-host run the only thing that is allowed to change my mind. A green laptop is a hint, not a signature I would paste into a field note as proof. If the clean run fails, I fix the helper or I delete the claim, and I do not split the difference. If a free MonkeyCode server is already in your scratch setup, run this pair of files there next.
Top comments (0)