DEV Community

ToNBi
ToNBi

Posted on

When a reasoning model's <think> block leaks a JSON draft into your parser

Disclosure: I wrote this article together with an AI assistant (Claude). The AI ran all the code below and checked the results. I am not a professional programmer, so please read it with that in mind.

The problem

When you pull JSON out of an LLM's reply, json.loads raising an error is the easy case. The harder case is when nothing fails and the wrong JSON quietly comes back.

Some reasoning models print their thinking before the answer, wrapped in tags like <think>...</think>. Sometimes that thinking contains a draft of the JSON:

<think>
The user wants {"name": "...", "age": ...}. Example: {"x": 1}
</think>
{"name": "Ann", "age": 30}
Enter fullscreen mode Exit fullscreen mode

(Depending on the model or provider, the thinking may come in a separate field instead. This article is about the case where it arrives mixed into the text.)

Attempt 1: first "{" to last "}"

text[text.find("{"): text.rfind("}") + 1]
Enter fullscreen mode Exit fullscreen mode

This gave me a JSONDecodeError, because the draft and the answer get cut out as one piece. At least it fails loudly.

Attempt 2: a repair library

I passed the same text to json-repair (version 0.63.5):

from json_repair import repair_json
repair_json(text, return_objects=True)
Enter fullscreen mode Exit fullscreen mode
[{'name': '...', 'age': '...'}, {'x': 1}, {'name': 'Ann', 'age': 30}]
Enter fullscreen mode Exit fullscreen mode

No error. All three JSON objects come back as a list. If you do not validate against a schema, this slips through.

It gets worse when the output is cut off

If the output stops in the middle of the thinking (for example, at a token limit):

<think>Hmm, maybe {"name": "Ann"}
Enter fullscreen mode Exit fullscreen mode

the same library returns:

{'name': 'Ann'}
Enter fullscreen mode Exit fullscreen mode

An unfinished draft is returned as if it were the final answer. The value is often "almost right", so it is easy to miss.

You can reproduce this in a few seconds:

pip install json-repair
python -c "from json_repair import repair_json; print(repair_json('<think>Hmm, {\"name\": \"Ann\"} maybe', return_objects=True))"
Enter fullscreen mode Exit fullscreen mode

The fix: throw the thinking away first

Two steps:

  1. Remove the <think> block before looking for JSON (including an unclosed one).
  2. From what is left, take the last complete JSON value.

Standard library only:

import json
import re

THINK_BLOCK = re.compile(r"<(think|thinking|reasoning)\b[^>]*>.*?</\1\s*>", re.S | re.I)
THINK_OPEN = re.compile(r"<(think|thinking|reasoning)\b[^>]*>.*\Z", re.S | re.I)


def extract_json(text: str):
    """Return the last complete JSON value found, or None."""
    text = THINK_BLOCK.sub("", text)   # drop closed thinking blocks
    text = THINK_OPEN.sub("", text)    # drop an unclosed (cut-off) one
    dec = json.JSONDecoder()
    found, i = None, 0
    while i < len(text):
        if text[i] in "{[":
            try:
                found, end = dec.raw_decode(text, i)
                i = end                # skip what was read (no nested picks)
                continue
            except json.JSONDecodeError:
                pass
        i += 1
    return found
Enter fullscreen mode Exit fullscreen mode

raw_decode tells you whether JSON can be read from a given position, and how far it reaches. That is why chatter before or after the JSON does not matter.

What I checked

Input Result
Draft JSON inside the thinking only the answer
Text before and after the JSON only the JSON
JSON inside a Markdown code fence only the JSON
Nested JSON the whole outer object
Truncated JSON None
Unclosed thinking block None

The last two matter most: if there is no answer yet, return None instead of guessing. That lets the caller treat it as a failure.

What this does not fix

Trailing commas, single quotes, and truncated JSON like {"a": 1 still need a repair library. But order matters:

text = THINK_BLOCK.sub("", text)
text = THINK_OPEN.sub("", text)
data = repair_json(text, return_objects=True)
Enter fullscreen mode Exit fullscreen mode

Drop the thinking, then repair the syntax, then validate against a schema.

Design notes

  • Never report a failed repair as a success. Return None or raise.
  • Log what you removed (for example, "removed 1 thinking block").
  • Validate with a schema as the last line of defense.

I am not an expert in this area. The code and results come from actually running them with an AI assistant, and output formats differ between models and providers, so your results may differ. If you know other failure modes or better approaches, I would be glad to hear them in the comments.

Top comments (0)