DEV Community

Cover image for Building an Autonomous CI/CD Log Triage & Self-Healing PR Engine with Python and DuckDB
Ruesch Manny
Ruesch Manny

Posted on Originally published at mannyverse767.gumroad.com

Building an Autonomous CI/CD Log Triage & Self-Healing PR Engine with Python and DuckDB

Building an Autonomous CI/CD Log Triage & Self-Healing PR Engine with Python and DuckDB

Every growing engineering team hits the same wall: CI pipelines fail, developers ignore cryptic 200MB console outputs, and senior engineers spend 15+ hours every week parsing raw text to fix the same 4 recurring issues—broken lockfiles, misconfigured test fixtures, missing environment variables, or dependency drift.

Observability vendors sell multi-thousand-dollar enterprise suites to ingest these logs, yet they stop at alerting. You still have to open the console, copy the stack trace, diagnose the problem, create a branch, write the fix, and submit a PR.

In this guide, we'll build a zero-infrastructure, self-healing pipeline that runs natively inside your CI environment (GitHub Actions, GitLab CI, or n8n runner). It parses multi-gigabyte build artifacts using DuckDB, clusters stack traces deterministically without bloated vector stores, prompts an LLM with compact context, and automatically submits self-healing Pull Requests via the GitHub REST API.


1. The Bottleneck: Why Traditional Log Triage Fails

Most CI log monitoring tools fail for three architectural reasons:

  1. Data Transfer Tax & Latency: Pushing gigabytes of raw container logs over the wire to external aggregators creates latency and ballooning bandwidth bills.
  2. Vector DB Overhead: Vector embeddings for error logs are overkill. Logs are semi-structured text where 95% of lines are identical runtime noise. Embedding entire raw logs produces polluted semantic spaces that fail to isolate exact trace boundaries.
  3. Lack of Actionability: Alerting on Slack is not a resolution. If a failure stems from a deprecated dependency version or a missing mock key, the remediation step is deterministic.

By executing triage locally on the CI runner, we eliminate network egress, process logs in memory via columnar storage, and resolve routine failures without human intervention.


2. System Architecture

The engine follows a strict four-stage pipeline running directly on your CI execution container:

[ Raw CI Log Files (*.log, *.json) ]
              │
              ▼
┌───────────────────────────────┐
│ DuckDB High-Throughput Ingest │ ──► Columnar filtering of timestamps,
│ & Regex Pattern Clustering    │     log levels, and stack frames
└──────────────┬────────────────┘
              │
              ▼
┌───────────────────────────────┐
│ Deterministic Triage Engine   │ ──► Isolates exact stack trace and
│ (Root-Cause Classifier)       │     queries target source file
└──────────────┬────────────────┘
              │
              ▼
┌───────────────────────────────┐
│ Structured LLM Patch Synthesizer │ ──► Generates unified diff
└──────────────┬────────────────┘
              │
              ▼
[ GitHub REST API: Creates Git Branch + Opens Self-Healing PR ]
Enter fullscreen mode Exit fullscreen mode
  • DuckDB: Reads gigabyte-scale logs directly from disk or gzipped archives using zero-copy SQL queries.
  • Pattern Clustering: Groups repeated log lines using tokenized signature matching.
  • Autonomous Git Engine: Validates the proposed diff locally, branches out, commits the patch, and triggers a PR.

3. The Code & Implementation

Step 1: Sub-Second Log Parsing with DuckDB

Instead of loading logs line-by-line in pure Python, we leverage DuckDB's native regex parsing engine to query the raw log file as a relational table.

import duckdb
import os

def extract_critical_failures(log_path: str) -> list[dict]:
    """
    Parses uncompressed or gzipped logs and extracts error frames
    without loading the entire file into Python memory.
    """
    con = duckdb.connect(database=":memory:")

    query = """
    WITH raw_lines AS (
        SELECT 
            row_number() OVER () AS line_num,
            column0 AS content
        FROM read_csv(?, header=False, sep='\n', auto_detect=False)
    ),
    filtered_errors AS (
        SELECT 
            line_num, 
            content,
            regexp_extract(content, '(?i)(error|fatal|exception|failed):?\s*(.*)', 2) AS error_msg
        FROM raw_lines
        WHERE regexp_matches(content, '(?i)(ERROR|FATAL|Exception|FAILED)')
    )
    SELECT 
        line_num, 
        content, 
        error_msg
    FROM filtered_errors
    LIMIT 25;
    """

    results = con.execute(query, [log_path]).fetchall()
    con.close()

    return [{"line": row[0], "raw": row[1], "error": row[2]} for row in results]
Enter fullscreen mode Exit fullscreen mode

Step 2: Extracting Context Around the Root Failure

Once DuckDB pinpoints the fatal line number, we extract the local stack trace window and the affected source code file to pass to our LLM generator.

import re
from pathlib import Path

def isolate_trace_context(log_path: str, error_line: int, window: int = 15) -> str:
    """Extracts preceding and succeeding context around an error line."""
    start = max(1, error_line - window)
    end = error_line + window

    con = duckdb.connect(database=":memory:")
    query = """
    SELECT list(column0)
    FROM (
        SELECT row_number() OVER () AS rn, column0
        FROM read_csv(?, header=False, sep='\n', auto_detect=False)
    )
    WHERE rn BETWEEN ? AND ?;
    """
    lines = con.execute(query, [log_path, start, end]).fetchone()[0]
    con.close()
    return "\n".join(lines) if lines else ""

def locate_offending_file(trace_text: str) -> tuple[str, str]:
    """Locates local source code file and extracts code content."""
    file_match = re.search(r'File "([a-zA-Z0-9_\-\./]+\.py)", line (\d+)', trace_text)
    if not file_match:
        return "", ""

    file_path, line_no = file_match.group(1), int(file_match.group(2))
    if Path(file_path).exists():
        return file_path, Path(file_path).read_text()
    return "", ""
Enter fullscreen mode Exit fullscreen mode

Step 3: LLM Patch Synthesizer & GitHub PR Automation

We feed the trace and target source code into a structured LLM prompt that enforces raw unified diff output, then push the patch directly through the GitHub API.

import json
import urllib.request
import base64

def generate_patch_with_llm(trace: str, source_code: str, file_path: str, api_key: str) -> str:
    """Generates a strict unified diff to fix the identified CI breakage."""
    system_prompt = (
        "You are an automated CI remediation agent. Analyze the provided test failure "
        "and source code. Output ONLY a valid unified diff (git patch) that fixes the failure. "
        "Do not output markdown code blocks, explanations, or commentary."
    )
    user_prompt = f"TARGET FILE: {file_path}\n\nSTACK TRACE:\n{trace}\n\nSOURCE CODE:\n{source_code}"

    payload = json.dumps({
        "model": "gpt-4o-mini",
        "messages": [
            {"role": "system", "content": system_prompt},
            {"role": "user", "content": user_prompt}
        ],
        "temperature": 0.1
    }).encode("utf-8")

    req = urllib.request.Request(
        "https://api.openai.com/v1/chat/completions",
        data=payload,
        headers={"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
    )

    with urllib.request.urlopen(req) as resp:
        data = json.loads(resp.read().decode("utf-8"))
        return data["choices"][0]["message"]["content"].strip()

def create_self_healing_pr(repo: str, branch: str, file_path: str, updated_content: str, token: str):
    """Applies remediation by opening a Git branch and PR via GitHub REST API."""
    headers = {
        "Authorization": f"Bearer {token}",
        "Accept": "application/vnd.github.v3+json"
    }
    base_url = f"https://api.github.com/repos/{repo}"

    # 1. Fetch default branch SHA
    req = urllib.request.Request(f"{base_url}/git/ref/heads/main", headers=headers)
    with urllib.request.urlopen(req) as resp:
        sha = json.loads(resp.read().decode())["object"]["sha"]

    # 2. Create target branch
    branch_payload = json.dumps({"ref": f"refs/heads/{branch}", "sha": sha}).encode()
    req = urllib.request.Request(f"{base_url}/git/refs", data=branch_payload, headers=headers)
    urllib.request.urlopen(req)

    # 3. Fetch file SHA
    req = urllib.request.Request(f"{base_url}/contents/{file_path}?ref={branch}", headers=headers)
    with urllib.request.urlopen(req) as resp:
        file_sha = json.loads(resp.read().decode())["sha"]

    # 4. Commit updated file
    encoded_content = base64.b64encode(updated_content.encode()).decode()
    commit_payload = json.dumps({
        "message": "fix(ci): autonomous self-healing patch for test failure",
        "content": encoded_content,
        "sha": file_sha,
        "branch": branch
    }).encode()
    req = urllib.request.Request(f"{base_url}/contents/{file_path}", data=commit_payload, headers=headers, method="PUT")
    urllib.request.urlopen(req)

    # 5. Open Pull Request
    pr_payload = json.dumps({
        "title": "🤖 [Self-Healing] Fix CI Failure Regression",
        "head": branch,
        "base": "main",
        "body": "This Pull Request was automatically generated by the Autonomous CI Log Triage Engine after detecting deterministic failures in the pipeline."
    }).encode()
    req = urllib.request.Request(f"{base_url}/pulls", data=pr_payload, headers=headers)
    urllib.request.urlopen(req)
Enter fullscreen mode Exit fullscreen mode

4. Deployment, Timeouts, and Rate Limits

To run this in production without adding latency to green runs:

  1. Conditional Triggers: Only execute the triage step on if: failure() in GitHub Actions or when: on_failure in GitLab CI.
  2. Run-Time Budgets: Restrict the DuckDB scan window to the last 5,000 lines. DuckDB handles this in under 200ms.
  3. LLM Cost Guardrails: Strip ANSI escape codes (\x1b\[[0-9;]*m) before sending payloads to LLMs to prevent token inflation. Truncate context to include only the stack trace boundary and the active function block.
  4. Prevent PR Flooding: Cache the hash of the error signature (MD5(trace_text)) in your build cache. If an identical PR is already pending, skip recreation.

5. Conclusion & Ready-to-Use Workflow

You can implement this architecture using the code snippets above or integrate it into an automated n8n webhook listener. By replacing third-party SaaS log aggregators with embedded DuckDB queries, you gain deterministic triage while keeping logs completely inside your security perimeter.

If you prefer a plug-and-play solution, you can get the complete turnkey package containing production GitHub Action workflows, Docker container configurations, synthetic log test fixtures, and pre-built n8n triage workflows:

Top comments (0)