DEV Community

PhenoX
PhenoX

Posted on

Asynchronous Parallel Validation & Diff Report Generation Tool for Multiple AI Platforms

Here is the fully translated and refined English article, tailored for a technical audience on Dev.to. All technical details, structural elements, and the complete code block have been preserved and translated into professional engineering terminology. The conclusion has been distilled to focus purely on the technical takeaways.


Asynchronous Parallel Validation & Diff Report Generation Tool for Multiple AI Platforms

Architecture Design and Implementation Choices

To achieve a script that is highly portable, runs immediately in any environment without prior setup, and minimizes dependencies, we adopted the following architectural design:

  1. Zero Third-Party Dependencies: The entire implementation relies solely on the Python standard library, specifically urllib.request for HTTP communications, alongside json and difflib.
  2. Multi-Process Parallelization: To circumvent the limitations of Python's Global Interpreter Lock (GIL) and enhance the parallel processing efficiency of CPU-bound tasks and multiple I/O operations, we utilized ProcessPoolExecutor.
  3. Global Timeout Control: To strictly constrain the total execution time to under 10 seconds, the remaining time is dynamically calculated within the as_completed loop, applying limits via future.result(timeout=remaining).

Below is the complete code that was evaluated during our testing phase—a script that ultimately succumbed to deadlocks and cascading timeouts.

import sys
import json
import urllib.request
import urllib.error
from concurrent.futures import ProcessPoolExecutor, as_completed
import time
import difflib

def call_endpoint(endpoint_info):
    """
    Sends an HTTP request to a single endpoint and returns the result.
    Placed at the module top-level to allow serialization by the process pool.
    """
    name = endpoint_info.get("name", "Unknown")
    url = endpoint_info.get("url")
    headers = endpoint_info.get("headers", {})
    payload = endpoint_info.get("payload", {})
    timeout = endpoint_info.get("timeout", 10)

    start_time = time.time()
    try:
        data = json.dumps(payload).encode("utf-8")
        req = urllib.request.Request(url, data=data, headers=headers, method="POST")

        with urllib.request.urlopen(req, timeout=timeout) as response:
            elapsed = time.time() - start_time
            body = response.read().decode("utf-8")
            try:
                parsed_body = json.loads(body)
            except json.JSONDecodeError:
                parsed_body = body

            return {
                "name": name,
                "status": "success",
                "status_code": response.status,
                "elapsed": round(elapsed, 3),
                "response": parsed_body
            }
    except urllib.error.HTTPError as e:
        elapsed = time.time() - start_time
        err_body = e.read().decode("utf-8", errors="ignore")
        return {
            "name": name,
            "status": "http_error",
            "status_code": e.code,
            "elapsed": round(elapsed, 3),
            "error": err_body
        }
    except urllib.error.URLError as e:
        elapsed = time.time() - start_time
        return {
            "name": name,
            "status": "url_error",
            "status_code": None,
            "elapsed": round(elapsed, 3),
            "error": str(e.reason)
        }
    except Exception as e:
        elapsed = time.time() - start_time
        return {
            "name": name,
            "status": "timeout_or_unknown",
            "status_code": None,
            "elapsed": round(elapsed, 3),
            "error": str(e)
        }

def generate_markdown_report(results, test_case_name):
    report = []
    report.append(f"# AI Validation & Diff Report: {test_case_name}")
    report.append("\n## 1. Execution Summary\n")
    report.append(
        "| Endpoint | Status | HTTP Code | Latency (s) | Error Details |\n"
        "| :--- | :--- | :--- | :--- | :--- |"
    )

    success_responses = {}

    for r in results:
        name = r["name"]
        status = r["status"]
        code = r["status_code"] if r["status_code"] is not None else "-"
        elapsed = r["elapsed"]
        err = "-"

        if status == "success":
            success_responses[name] = r["response"]
        else:
            err = r.get("error", "Unknown error").replace("\n", " ")

        report.append(f"| {name} | {status} | {code} | {elapsed} | {err} |")

    report.append("\n## 2. Response Outputs\n")
    for r in results:
        report.append(f"### [{r['name']}] Output")
        report.append("```

json")
        report.append(json.dumps(r.get("response") or r.get("error"), ensure_ascii=False, indent=2))
        report.append("

```\n")

    report.append("## 3. Structural & Textual Diff Analysis\n")
    names = list(success_responses.keys())
    if len(names) < 2:
        report.append("Skipping diff analysis because fewer than 2 successful responses were received.\n")
    else:
        for i in range(len(names)):
            for j in range(i + 1, len(names)):
                n1, n2 = names[i], names[j]
                text1 = json.dumps(success_responses[n1], ensure_ascii=False, indent=2).splitlines()
                text2 = json.dumps(success_responses[n2], ensure_ascii=False, indent=2).splitlines()

                diff = list(difflib.unified_diff(text1, text2, fromfile=n1, tofile=n2, lineterm=""))
                report.append(f"### Diff: {n1} vs {n2}")
                if diff:
                    report.append("```

diff")
                    report.extend(diff[:50])
                    if len(diff) > 50:
                        report.append("... (diff truncated)")
                    report.append("

```\n")
                else:
                    report.append("No differences found (Exact match).\n")

    return "\n".join(report)

def main():
    input_data = sys.stdin.read()
    if not input_data.strip():
        print("Error: Test cases must be provided via standard input.", file=sys.stderr)
        sys.exit(1)

    try:
        test_suite = json.loads(input_data)
    except json.JSONDecodeError as e:
        print(f"Error: Failed to parse JSON: {e}", file=sys.stderr)
        sys.exit(1)

    test_name = test_suite.get("test_name", "Unnamed Test")
    endpoints = test_suite.get("endpoints", [])

    if not endpoints:
        print("Error: No target endpoints defined in the test suite.", file=sys.stderr)
        sys.exit(1)

    results = []
    max_workers = min(len(endpoints), 16)

    global_timeout = 10.0
    start_time = time.time()

    with ProcessPoolExecutor(max_workers=max_workers) as executor:
        future_to_endpoint = {executor.submit(call_endpoint, ep): ep for ep in endpoints}

        for future in as_completed(future_to_endpoint):
            ep = future_to_endpoint[future]
            name = ep.get("name", "Unknown")

            elapsed_total = time.time() - start_time
            remaining = max(0.1, global_timeout - elapsed_total)

            try:
                res = future.result(timeout=remaining)
                results.append(res)
            except Exception as e:
                results.append({
                    "name": name,
                    "status": "timeout_or_error",
                    "status_code": None,
                    "elapsed": round(time.time() - start_time, 3),
                    "error": str(e) or "Execution timed out or failed"
                })

    markdown_report = generate_markdown_report(results, test_name)
    print(markdown_report)

if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

đź’ˇ For immediate deployment: The complete source code suite (ZIP) for this architecture is available on Gumroad for $0+ (Pay What You Want).


The Quagmire: Why the "10-Second Barrier" Remained Unbroken

After repeating our QA test suite three times, the precise mechanisms that dragged this system into fatal deadlocks and latency traps became evident.

1. The Mismatch of ProcessPoolExecutor for I/O-Bound Workloads

While multiprocessing is highly effective for CPU-intensive tasks, it is fundamentally unsuitable for network I/O-heavy operations like calling external LLM APIs.
The overhead introduced by Inter-Process Communication (IPC), object serialization/deserialization, and OS-level process spawning is not negligible. When multiple external APIs simultaneously experienced delayed responses, the aggressive context switching between processes severely bottlenecked system resources.

2. The Collapse of Dynamic future.result(timeout=remaining) Control

To enforce the strict 10-second global limitation, we integrated the following calculation into the core loop:

elapsed_total = time.time() - start_time
remaining = max(0.1, global_timeout - elapsed_total)
res = future.result(timeout=remaining)
Enter fullscreen mode Exit fullscreen mode

The fatal flaw here is that as_completed yields futures in the order they complete.
If the interpreter first evaluates a future from an endpoint that experienced heavy latency (e.g., a local LLM taking 8 seconds to respond), the remaining time shrinks drastically. Consequently, when the loop processes the next future—even one from an API that responded almost instantaneously—it applies the severely depleted remaining time (e.g., 0.1 seconds). This triggers a cascading timeout collapse, forcefully throwing TimeoutExpired exceptions and terminating requests that actually succeeded at the network layer.

3. Misjudging the Trade-offs Against Asynchronous I/O (asyncio)

Driven by an absolute constraint to maintain "zero third-party dependencies," we attempted to forcefully parallelize the blocking urllib.request using multi-processing, rather than adopting Python's native asyncio combined with a non-blocking wrapper (like http.client wrapped in an executor) or leveraging asynchronous runtimes.
By abandoning efficient event-loop multiplexing (which threads or async coroutines handle natively), we constructed a fragile, synchronous wait-state architecture wholly dependent on process scaling—a fundamentally broken approach for scaling HTTP concurrency.


Architectural Insights and Anti-Patterns Learned

Based on the failures observed in this prototype, we offer the following critical insights for engineers designing similar asynchronous validation pipelines:

  • Never Use ProcessPoolExecutor for I/O-Bound API Multiplexing For parallel network requests, event-loop-based asynchronous I/O (such as asyncio coupled with httpx or native asyncio sockets) is the strictly superior choice. The overhead of multi-processing is excessively punitive for anything outside of purely CPU-bound domains.
  • Decouple "Global Timeouts" from "Per-Task Timeouts" When enforcing a global execution time limit, dynamically slicing individual task wait times inside a consumer loop creates brittle, order-dependent bugs. Each request should maintain its own independent timeout threshold. Global limits must be orchestrated via cancellable context managers or execution wrappers (e.g., utilizing the timeout parameter in asyncio.wait()), ensuring deterministic behavior regardless of the order of completion.
  • Evaluate the True Cost of "Zero Third-Party Dependencies" Attempting to write highly robust asynchronous HTTP handlers and structured JSON diffing using strictly the Python standard library inevitably leads to reinventing the wheel. This artificially inflates code complexity and maintenance debt. Depending on your operational constraints, integrating lightweight, reliable third-party libraries (like httpx) is often a significantly safer and more scalable engineering decision.

Conclusion

A design that appeared conceptually sound on paper was ultimately shattered against the 10-second barrier by real-world network jitter and a fundamental misapplication of the process concurrency model. However, highlighting the exact boundary limits of standard library concurrency models serves as a solid foundation for more robust architectural decisions in the future.

In software engineering, the structural understanding of why failing code breaks is the most valuable asset for future success. We hope this postmortem acts as a navigational compass for developers stepping into the demanding territory of highly concurrent API validation.


If this engineering log saved your production server (and your sanity), consider supporting our architecture on GitHub Sponsors.
Sponsor on GitHub

Top comments (0)