If you've shipped anything on top of an LLM, you know this pain: you ask the model for JSON, and you get back JSON wrapped in a code fence, with a friendly "Sure! Here's your data:" in front, single quotes instead of double, a trailing comma, and — on a bad day — the last brace missing because the response got cut off.
So you write a regex. Then another. Then a try/except that strips fences. Then you discover the model sometimes emits True instead of true. Six months later your "JSON cleaner" is 200 lines of defensive string surgery that everyone is afraid to touch.
I got tired of copy-pasting that file between projects, so I turned it into a small API. Here's what it handles and how it works.
The failure modes
Real LLM output that should be JSON tends to break in a predictable handful of ways:
- Markdown code fences around the JSON.
- Prose around it — "Here is the result: { ... } Hope this helps!"
- Single quotes instead of double.
- Unquoted keys.
- Trailing commas.
- Python/JS literals — None, True, False, NaN, undefined.
- Inline comments.
- Truncation — the response hit the token limit and just... stops.
The first seven are cosmetic. The eighth is the nasty one: you have to walk the string, track the open braces and brackets and whether you're inside a string, and close everything back up in the right order.
One call
Send the broken string; get back valid parsed JSON, a normalized text version, and the list of fixes that were applied. The truncated case is closed automatically — a fragment that starts an object with an unclosed array comes back as a complete, valid object.
Why deterministic, not "ask another model to fix it"
You could send the broken JSON to a second LLM call and ask it to fix it. But that's slower, costs tokens, is non-deterministic, and can hallucinate values that were never there. Repairing JSON is a parsing problem, not a reasoning problem — so it should be solved with a parser. Same input, same output, every time, in milliseconds.
It's part of a small toolbox
The same API also has the endpoints I kept needing next to JSON repair:
- context-compress — trim text to a token budget before sending it to a model (removes HTML and boilerplate, dedupes, cuts on sentence boundaries).
- srt-clean — turn SRT/VTT captions into clean text for prompting, and tell you how many tokens you saved.
- log-triage — paste raw Docker/systemd/stack-trace logs and get the exception type, message, and likely cause.
If this saves you from maintaining yet another clean_json.py, it's live on RapidAPI with a free tier: https://rapidapi.com/cesaricf79/api/llm-dev-utilities
What's your worst "the model returned almost-JSON" story? I'm curious how many of these failure modes I'm still missing.
Top comments (0)