DEV Community

PhenoX
PhenoX

Posted on

Building a CLI Validator to Forcefully Reconcile Unstable LLM Outputs: A Single-File JSON Schema Alignment Tool

Title: Building a CLI Validator to Forcefully Reconcile Unstable LLM Outputs: A Single-File JSON Schema Alignment Tool

Why Existing Validators Just Don't Cut It

If you throw raw LLM outputs directly into json.loads or jsonschema, they will shatter in seconds. The reality of production LLM pipelines is full of chaos like this:

  1. Markdown Code Block Invasions: Out of pure "helpfulness," LLMs often wrap their JSON outputs in json ` or `.
  2. Chronic Syntax Errors: Trailing commas and stray single quotes are everywhere.
  3. Schema Deviations: Required properties randomly drop out, or type mismatches occur.

Writing dozens of lines of ad-hoc parsing functions to prevent this is a nightmare. Moreover, if you write naive regular expressions that get caught in infinite loops when handling massive, malformed payloads, your entire pipeline will hang.

To solve this, I designed a single-file CLI tool that meets the following stringent requirements:

  • Sub-second One-Shot Execution: Fast enough to be embedded at the front of CI/CD pipelines or nightly batches.
  • 1MB Input Size Limit: Prevents memory exhaustion and infinite regex loops (ReDoS).
  • Robust Fallback Mechanisms: Automatic Markdown stripping, trailing comma removal, and quote replacement.
  • Strict Validation & Default Value Imputation via JSON Schema

The Completed Alignment Tool

The only dependencies are the standard library and jsonschema (optional). It is implemented as a drop-in, single-file Python 3 script that can be executed instantly anywhere.

#!/usr/bin/env python3
"""
JSON Schema Alignment CLI Utility (TOAI2 Custom Edition - Timeout Safe)
A one-shot CLI tool that forcefully reconciles fluctuating LLM outputs and strictly enforces a given JSON Schema.
"""

import sys
import json
import re
import argparse
from typing import Any, Dict, Optional, Tuple

try:
    import jsonschema
except ImportError:
    jsonschema = None

MAX_INPUT_SIZE = 1024 * 1024  # 1MB limit to prevent infinite hangs and memory exhaustion


def clean_and_parse_json(raw_text: str) -> Any:
    """
    Strips Markdown blocks and common syntax errors (e.g., trailing commas) from malformed JSON text,
    and forcefully attempts to parse it (timeout and infinite loop safe).
    """
    if len(raw_text) > MAX_INPUT_SIZE:
        raise ValueError("Input data size is too large (1MB limit).")

    text = raw_text.strip()

    # 1. Strip Markdown code blocks (safe non-greedy match)
    match = re.search(r"```

(?:json)?\s*([\s\S]*?)\s*

```", text)
    if match:
        text = match.group(1).strip()
    else:
        text = re.sub(r"^```

(?:json)?", "", text)
        text = re.sub(r"

```$", "", text)
        text = text.strip()

    # 2. Safe regex replacement for common LLM output corruption (e.g., trailing commas)
    text = re.sub(r",\s*([\}\]])", r"\1", text)

    # 3. Attempt standard JSON parsing
    try:
        return json.loads(text)
    except json.JSONDecodeError as e:
        try:
            # Simple fallback: replace single quotes with double quotes
            fixed_text = text.replace("'", '"')
            return json.loads(fixed_text)
        except json.JSONDecodeError:
            raise ValueError(f"Failed to parse JSON: {e}. Input snippet: {text[:100]}...")


def align_to_schema(data: Any, schema: Optional[Dict[str, Any]]) -> Tuple[Any, list]:
    """
    Validates data against a JSON Schema and imputes default values where necessary.
    """
    errors = []
    if schema is None:
        return data, errors

    if jsonschema is None:
        raise ImportError(
            "Schema validation was requested, but the 'jsonschema' module is not installed."
        )

    try:
        validator = jsonschema.Draft202012Validator(schema)
        if isinstance(data, dict) and "properties" in schema:
            for prop, prop_schema in schema.get("properties", {}).items():
                if prop not in data and "default" in prop_schema:
                    data[prop] = prop_schema["default"]

        validator.validate(data)
    except jsonschema.ValidationError as ve:
        errors.append(str(ve))
    except Exception as ex:
        errors.append(f"Schema validation error: {str(ex)}")

    return data, errors


def main():
    parser = argparse.ArgumentParser(description="LLM Output JSON Forced Repair & Alignment Tool")
    parser.add_argument("-f", "--file", help="Path to input JSON file (uses stdin if omitted)", default=None)
    parser.add_argument("-s", "--schema", help="Path to JSON Schema file for validation", default=None)
    args = parser.parse_args()

    # Retrieve input (safe stream reading)
    if args.file:
        try:
            with open(args.file, "r", encoding="utf-8") as f:
                raw_input = f.read(MAX_INPUT_SIZE + 1)
        except Exception as e:
            print(json.dumps({"error": f"File read error: {str(e)}"}, ensure_ascii=False), file=sys.stderr)
            sys.exit(1)
    else:
        raw_input = sys.stdin.read(MAX_INPUT_SIZE + 1)

    if len(raw_input) > MAX_INPUT_SIZE:
        print(json.dumps({"error": "Input data exceeds allowed size."}, ensure_ascii=False), file=sys.stderr)
        sys.exit(1)

    if not raw_input.strip():
        print(json.dumps({"error": "Input is empty."}, ensure_ascii=False), file=sys.stderr)
        sys.exit(1)

    # Load schema
    schema_data = None
    if args.schema:
        try:
            with open(args.schema, "r", encoding="utf-8") as f:
                schema_data = json.load(f)
        except Exception as e:
            print(json.dumps({"error": f"Schema file read error: {str(e)}"}, ensure_ascii=False), file=sys.stderr)
            sys.exit(1)

    # Execute processing (timeout & exception safe)
    try:
        parsed_data = clean_and_parse_json(raw_input)
        aligned_data, validation_errors = align_to_schema(parsed_data, schema_data)

        output = {
            "status": "success",
            "validation_errors": validation_errors,
            "data": aligned_data
        }
        print(json.dumps(output, ensure_ascii=False, indent=2))

    except Exception as e:
        error_response = {
            "status": "error",
            "message": str(e),
            "raw_input_snippet": raw_input[:200]
        }
        print(json.dumps(error_response, ensure_ascii=False, indent=2), file=sys.stderr)
        sys.exit(1)


if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

💡 For immediate deployment: The complete source code suite (ZIP) for this architecture is available on Gumroad for $0+ (Pay What You Want).


Landmines and Gritty Workarounds from the Trenches

Before putting this tool into production, I learned a few painful lessons.

1. The Terror of Infinite Hangs from Greedy Matching

Initially, I used a sloppy greedy match r".*" to strip Markdown code blocks. When an LLM went off the rails and output a massive wall of text without a closing tag, the regex engine drowned in a vortex of catastrophic backtracking. The CPU usage pinned at 100%, and the process completely hung.
The Fix: I switched to a safe, non-greedy match using r"(?:json)?\s*([\s\S]*?)\s*". More importantly, by enforcing a physical 1MB size limit (MAX_INPUT_SIZE) upstream, I completely eliminated the risk of resource exhaustion at the hardware level.

2. Enforcing "Structured Output" During Fatal Errors

If the script crashes with an exception and dumps a raw Python traceback (Traceback...) to stderr, downstream pipelines (like shell scripts or Go processes) will fail to parse it, causing cascading crashes.
The Fix: Even during abnormal terminations, the tool is designed to always emit a structured JSON payload (status: "error" and a snippet) to stderr. This standardizes error handling for upstream systems and prevents domino effects in automated pipelines.


How to Integrate It into Your Pipeline

Since this CLI fully supports standard input and output, it perfectly embodies the UNIX pipeline philosophy.

# Pipe raw LLM output to instantly perform schema validation and default value imputation
cat llm_raw_output.txt | python3 align_validator.py -s schema.json
Enter fullscreen mode Exit fullscreen mode

On success, it returns a pristine, structured JSON object like this:

{
  "status": "success",
  "validation_errors": [],
  "data": {
    "id": "uuid-1234",
    "status": "active",
    "score": 95
  }
}
Enter fullscreen mode Exit fullscreen mode

Conclusion

When integrating LLMs into production backends, we are constantly forced to act as interpreters between "probabilistic, capricious outputs" and "deterministic, cold-blooded databases."

Instead of relying solely on prompt engineering voodoo, inserting a "forceful reconciler" like this at the infrastructure boundary acts as a robust fail-safe. It effectively bridges the gap between raw generative AI and traditional software architectures, keeping your pipelines stable when the models inevitably hallucinate formatting errors.


If this engineering log saved your production server (and your sanity), consider supporting our architecture on GitHub Sponsors.
Sponsor on GitHub

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to