DEV Community

Daniel Dong
Daniel Dong

Posted on

Stop regex-parsing your LLM output. Make it return JSON.

The moment an LLM hands you a wall of prose and you reach for a regex, you've moved the problem — you traded a model that knows the answer for a string you have to re-parse to get it back.

curl https://aibridge-api.com/v1/chat/completions \
  -H "Authorization: Bearer mb-xxxxxxxx" \
  -d '{
    "model": "deepseek-v4-pro",
    "messages": [
      {"role": "system", "content": "You return ONLY valid JSON. No markdown fences, no prose, no commentary."},
      {"role": "user", "content": "Extract name, date, and amount from this invoice text: \"...\""}
    ]
  }'
Enter fullscreen mode Exit fullscreen mode

The response lands as clean, machine-readable JSON you can feed straight into your pipeline — no .split(), no regex, no "it worked until the model added an emoji."

The prose-parsing trap

LLMs default to prose because prose is how they were trained to talk. But your code doesn't want prose — it wants fields. So developers end up doing one of three fragile things:

  1. Regex surgery — brittle, breaks the moment the model rephrases
  2. String splitting — "find the first colon, take everything after it" (works until it doesn't)
  3. "Just ask nicely" — prompt the model to return JSON, but never verify it actually did

Every one of those fails in production, usually at 2 AM, usually on a customer's most important request.

The fix is a contract, not a wish

Structured output works when you treat the format as a requirement, not a request:

  • Say it in the system prompt — "return ONLY valid JSON, no fences, no prose"
  • Give it a schema — describe the exact keys and types you expect, or better, paste an example
  • Verify before you trust — parse the result with json.loads / JSON.parse; if it throws, retry with a harder prompt

The model is good at following a clear contract. It's bad at guessing what you meant. Spell out the shape and it will usually deliver it.

Where structured output pays for itself

  • Extraction — pull name/date/amount from messy text into clean fields
  • Classification — return {"label": "refund", "confidence": 0.97} instead of "This appears to be a refund request"
  • API responses — build an endpoint that returns LLM-shaped JSON, not prose you have to clean
  • Chained pipelines — one model's JSON output becomes the next model's typed input, no parsing in between The moment your output is typed and parseable, your LLM call becomes a normal function in your codebase — testable, composable, boring in the best way.

One more trick: it pairs with the model router

You don't need a flagship model to extract fields from an invoice. Route the cheap, fast models to the structured-output work and save the heavy reasoning for the hard stuff:

def extract(messages) -> dict:
    # glm-4-flash is fast, cheap, and plenty for JSON extraction
    resp = client.chat.completions.create(model="glm-4-flash", messages=messages)
    return json.loads(resp.choices[0].message.content)  # typed, or it raises
Enter fullscreen mode Exit fullscreen mode
  • Fast tier — glm-4-flash, deepseek-v4-flash, moonshot-v1-8k
  • Balanced tier — glm-4-air, deepseek-chat, qwen-plus
  • Flagship tier — deepseek-v4-pro, qwen3-235b-a22b, glm-4-plus
  • Deep reasoning — deepseek-reasoner, kimi-k3 (1M context) 15 models, 4 vendors, one OpenAI-compatible endpoint.

Pricing that keeps extraction cheap

  • Free tier: 500K tokens/month (weighted)
  • Pro: $9.90/month for 5M tokens
  • Top-ups: 1M / $2.99 · 5M / $9.90 · 20M / $29.90 (never expire)

The takeaway

An LLM that returns JSON is a function. An LLM that returns prose is a chore. Put the format in the contract and let your code trust the output.
Try structured output on all 15 models, free.

→ aibridge-api.com · support@aibridge-api.com

1

2

3

4

5

Top comments (0)