DEV Community

ToNBi
ToNBi

Posted on AI-assisted

When you "repair" truncated JSON, you stop noticing it was truncated

Disclosure: I wrote this article together with an AI assistant (Claude). The AI wrote and ran all the code below and checked the results. I'm not a professional programmer, so please read it with that in mind.

LLM APIs return HTTP 200 even when the output stops because it hit the token limit. Nothing looks like an error, so a JSON reply that was cut off halfway can sail straight into the next step of your pipeline.

A common fix is to put a repair library such as json-repair in front of the parser. It closes the brackets you were missing and hands you valid JSON, which is very handy.

But when the JSON was cut off, "repairing" it also hides the fact that it was cut off. That's obvious in hindsight, but I had no idea how often it actually bites, so I tried every possible cut and counted.

Here's the short version:

  • Passed through a repair library, almost everything (98.5%) came back as plausible-looking JSON
  • Even with a JSON Schema where every field is required, 137 of 853 truncations (16.1%) still passed validation
  • Adding one "I finished writing" marker at the end of the JSON brought that down to zero

How I tested it

I wrote five valid JSON outputs of the kind an LLM might return (the string values were in Japanese):

Sample Contents Characters
Order extraction order id, customer, 3 items, total, note 233
Tool-call arguments path, old string, new string, replace-all flag 126
Email classification label, confidence, 3 reasons 106
Text summary title, summary, 3 tags 135
Event extraction date, time, title × 4 258

Each one was cut at every position, from the first character to one character before the end. That gives 853 truncated strings in total. Each went through three steps:

  • A: plain json.loads
  • B: json-repair (0.63.5)
  • C: B, then JSON Schema validation (all fields required, plus types, ranges and formats)

By "silently damaged" I mean: no error, a value came back, but it differs from the original JSON.

Results

Sample Truncation points A: parsed B: silently damaged C: still passed validation
Order extraction 232 0 230 (99.1%) 11 (4.7%)
Tool-call arguments 125 0 124 (99.2%) 0 (0.0%)
Email classification 105 0 102 (97.1%) 51 (48.6%)
Text summary 134 0 131 (97.8%) 17 (12.7%)
Event extraction 257 0 253 (98.4%) 58 (22.6%)
Total 853 0 840 (98.5%) 137 (16.1%)

json.loads raised an error on all 853. That's not a bad thing: an error at least tells you something is wrong.

The repair library returned something for 98.5% of them. That isn't a flaw in the library. Closing brackets so the result is readable is exactly its job. It's just that doing the job well wipes out the evidence of the cut.

Schema validation catches a lot, but 137 got through. Looking at them, there were only two kinds of damage (some truncations had both):

Damage Count Example
Array lost elements 107 3 reasons became ["S"]
String cut short 91 a sentence became its first letter

When the cut lands inside an array, the repair library simply closes the array there, and the remaining elements look like they never existed. A required-field check only asks whether the field is present, so it can't notice.

The tool-call sample had zero because it has no array and ends with a boolean. A boolean cut halfway turns into a string like "fa", and the type check catches that.

When a number comes last, the amount changes

In my five samples, a number cut short was still caught, because required fields after it were missing. So I also tried a case where the number is the last field:

'{"product": "Mechanical keyboard", "price": 1'     -> price: 1     passes validation
'{"product": "Mechanical keyboard", "price": 128'   -> price: 128   passes validation
'{"product": "Mechanical keyboard", "price": 1280'  -> price: 1280  passes validation
Enter fullscreen mode Exit fullscreen mode

The real value was 12800. 1280 is a tenth of that, but it's a perfectly valid integer, so the schema has no complaint.

Try it yourself

Here is a small reproduction using the email-classification example. The strings are in English here, so the counts differ from the table above (longer strings mean more cut points inside strings).

import json
from json_repair import repair_json            # pip install json-repair jsonschema
from jsonschema import Draft202012Validator

answer = {
    "label": "spam",
    "confidence": 0.97,
    "reasons": [
        "Sender domain does not match the company named in the body",
        "Asks for an urgent bank transfer",
        "Contains a shortened URL",
    ],
}
schema = {
    "type": "object",
    "required": ["label", "confidence", "reasons"],   # every field is required
    "properties": {
        "label": {"type": "string", "enum": ["spam", "ham"]},
        "confidence": {"type": "number", "minimum": 0, "maximum": 1},
        "reasons": {"type": "array", "items": {"type": "string"}, "minItems": 1},
    },
}

text = json.dumps(answer, ensure_ascii=False)
validator = Draft202012Validator(schema)
slipped = 0
for k in range(1, len(text)):                  # cut the text at every position
    got = repair_json(text[:k], return_objects=True)
    if got != answer and not list(validator.iter_errors(got)):
        slipped += 1                           # passed validation, yet differs from the original
        if slipped <= 3:
            print(json.dumps(got, ensure_ascii=False))
print(f"{slipped} / {len(text) - 1} truncation points: damaged data passed schema validation")
Enter fullscreen mode Exit fullscreen mode

Output:

{"label": "spam", "confidence": 0.97, "reasons": ["S"]}
{"label": "spam", "confidence": 0.97, "reasons": ["Se"]}
{"label": "spam", "confidence": 0.97, "reasons": ["Sen"]}
121 / 175 truncation points: damaged data passed schema validation
Enter fullscreen mode Exit fullscreen mode

Fix 1: check why the model stopped

The most reliable fix is to read the "why did generation stop" field the API gives you. The name differs by provider:

Provider Where to look Value when cut off
Anthropic (Claude) stop_reason max_tokens
OpenAI (Chat Completions) finish_reason length
OpenAI (Responses API) status and incomplete_details.reason incomplete / max_output_tokens
Google (Gemini) finishReason MAX_TOKENS

If the output was cut off, don't repair it. Raise the limit and ask again.

That said, frameworks can hide this field, and there have been reports of Gemini not setting finishReason when the limit is reached. So it's nice to have a safety net on the JSON side too.

Fix 2: end the JSON with a completion marker

When output is cut off, the end is always what gets lost. You can use that. Put a "_complete": true field last, and make the schema reject anything but true. If the model didn't finish, the field is either missing or half-written (like "tr"), and validation stops it.

answer = {
    "label": "spam",
    "confidence": 0.97,
    "reasons": ["Sender domain does not match the company named in the body", "Asks for an urgent bank transfer", "Contains a shortened URL"],
    "_complete": True,                          # <- a marker written last
}
schema = {
    "type": "object",
    "required": ["label", "confidence", "reasons", "_complete"],
    "properties": {
        "label": {"type": "string", "enum": ["spam", "ham"]},
        "confidence": {"type": "number", "minimum": 0, "maximum": 1},
        "reasons": {"type": "array", "items": {"type": "string"}, "minItems": 1},
        "_complete": {"const": True},           # anything but true is rejected
    },
}
Enter fullscreen mode Exit fullscreen mode

I added the marker to all five samples and cut them the same way. Of 943 truncations, none passed validation (0 of 943). On the English snippet above it was 0 of 194.

One caveat: this assumes the model writes the fields in schema order. Structured-output modes usually do, but check on your setup, and consider telling the model in the prompt to write _complete last.

If you're building a repair tool

I ran the same 853 cases through my own repair API. It did emit a warning on every one ("the JSON was never closed"), but for 141 of them (16.5%) it still returned success: true.

A warning in free text forces the caller to parse a string to find out. If you're building a repair step yourself, expose the signal as its own field, like truncated: true, so the caller can decide whether to use the result or retry.

Wrapping up

Repair libraries are useful, but on truncated JSON they erase the traces of the cut. Required fields stop most cases, but arrays that lose elements and strings that stop early still get through, 16.1% in this experiment. Checking the stop reason and ending the JSON with a marker work well together.

One more caveat: these percentages assume that a cut is equally likely at every position. Real outputs differ in length and in where they get cut, so real-world numbers will differ. There were also only five samples, so treat the numbers as "this can happen at about this scale".

If you know of other failure modes or better fixes, I'd love to hear about them in the comments.

Top comments (0)