Building an Autonomous CI/CD Log Triage & Self-Healing PR Engine with Python and DuckDB
Every growing engineering team hits the same wall: CI pipelines fail, developers ignore cryptic 200MB console outputs, and senior engineers spend 15+ hours every week parsing raw text to fix the same 4 recurring issues—broken lockfiles, misconfigured test fixtures, missing environment variables, or dependency drift.
Observability vendors sell multi-thousand-dollar enterprise suites to ingest these logs, yet they stop at alerting. You still have to open the console, copy the stack trace, diagnose the problem, create a branch, write the fix, and submit a PR.
In this guide, we'll build a zero-infrastructure, self-healing pipeline that runs natively inside your CI environment (GitHub Actions, GitLab CI, or n8n runner). It parses multi-gigabyte build artifacts using DuckDB, clusters stack traces deterministically without bloated vector stores, prompts an LLM with compact context, and automatically submits self-healing Pull Requests via the GitHub REST API.
1. The Bottleneck: Why Traditional Log Triage Fails
Most CI log monitoring tools fail for three architectural reasons:
- Data Transfer Tax & Latency: Pushing gigabytes of raw container logs over the wire to external aggregators creates latency and ballooning bandwidth bills.
- Vector DB Overhead: Vector embeddings for error logs are overkill. Logs are semi-structured text where 95% of lines are identical runtime noise. Embedding entire raw logs produces polluted semantic spaces that fail to isolate exact trace boundaries.
- Lack of Actionability: Alerting on Slack is not a resolution. If a failure stems from a deprecated dependency version or a missing mock key, the remediation step is deterministic.
By executing triage locally on the CI runner, we eliminate network egress, process logs in memory via columnar storage, and resolve routine failures without human intervention.
2. System Architecture
The engine follows a strict four-stage pipeline running directly on your CI execution container:
[ Raw CI Log Files (*.log, *.json) ]
│
▼
┌───────────────────────────────┐
│ DuckDB High-Throughput Ingest │ ──► Columnar filtering of timestamps,
│ & Regex Pattern Clustering │ log levels, and stack frames
└──────────────┬────────────────┘
│
▼
┌───────────────────────────────┐
│ Deterministic Triage Engine │ ──► Isolates exact stack trace and
│ (Root-Cause Classifier) │ queries target source file
└──────────────┬────────────────┘
│
▼
┌───────────────────────────────┐
│ Structured LLM Patch Synthesizer │ ──► Generates unified diff
└──────────────┬────────────────┘
│
▼
[ GitHub REST API: Creates Git Branch + Opens Self-Healing PR ]
- DuckDB: Reads gigabyte-scale logs directly from disk or gzipped archives using zero-copy SQL queries.
- Pattern Clustering: Groups repeated log lines using tokenized signature matching.
- Autonomous Git Engine: Validates the proposed diff locally, branches out, commits the patch, and triggers a PR.
3. The Code & Implementation
Step 1: Sub-Second Log Parsing with DuckDB
Instead of loading logs line-by-line in pure Python, we leverage DuckDB's native regex parsing engine to query the raw log file as a relational table.
import duckdb
import os
def extract_critical_failures(log_path: str) -> list[dict]:
"""
Parses uncompressed or gzipped logs and extracts error frames
without loading the entire file into Python memory.
"""
con = duckdb.connect(database=":memory:")
query = """
WITH raw_lines AS (
SELECT
row_number() OVER () AS line_num,
column0 AS content
FROM read_csv(?, header=False, sep='\n', auto_detect=False)
),
filtered_errors AS (
SELECT
line_num,
content,
regexp_extract(content, '(?i)(error|fatal|exception|failed):?\s*(.*)', 2) AS error_msg
FROM raw_lines
WHERE regexp_matches(content, '(?i)(ERROR|FATAL|Exception|FAILED)')
)
SELECT
line_num,
content,
error_msg
FROM filtered_errors
LIMIT 25;
"""
results = con.execute(query, [log_path]).fetchall()
con.close()
return [{"line": row[0], "raw": row[1], "error": row[2]} for row in results]
Step 2: Extracting Context Around the Root Failure
Once DuckDB pinpoints the fatal line number, we extract the local stack trace window and the affected source code file to pass to our LLM generator.
import re
from pathlib import Path
def isolate_trace_context(log_path: str, error_line: int, window: int = 15) -> str:
"""Extracts preceding and succeeding context around an error line."""
start = max(1, error_line - window)
end = error_line + window
con = duckdb.connect(database=":memory:")
query = """
SELECT list(column0)
FROM (
SELECT row_number() OVER () AS rn, column0
FROM read_csv(?, header=False, sep='\n', auto_detect=False)
)
WHERE rn BETWEEN ? AND ?;
"""
lines = con.execute(query, [log_path, start, end]).fetchone()[0]
con.close()
return "\n".join(lines) if lines else ""
def locate_offending_file(trace_text: str) -> tuple[str, str]:
"""Locates local source code file and extracts code content."""
file_match = re.search(r'File "([a-zA-Z0-9_\-\./]+\.py)", line (\d+)', trace_text)
if not file_match:
return "", ""
file_path, line_no = file_match.group(1), int(file_match.group(2))
if Path(file_path).exists():
return file_path, Path(file_path).read_text()
return "", ""
Step 3: LLM Patch Synthesizer & GitHub PR Automation
We feed the trace and target source code into a structured LLM prompt that enforces raw unified diff output, then push the patch directly through the GitHub API.
import json
import urllib.request
import base64
def generate_patch_with_llm(trace: str, source_code: str, file_path: str, api_key: str) -> str:
"""Generates a strict unified diff to fix the identified CI breakage."""
system_prompt = (
"You are an automated CI remediation agent. Analyze the provided test failure "
"and source code. Output ONLY a valid unified diff (git patch) that fixes the failure. "
"Do not output markdown code blocks, explanations, or commentary."
)
user_prompt = f"TARGET FILE: {file_path}\n\nSTACK TRACE:\n{trace}\n\nSOURCE CODE:\n{source_code}"
payload = json.dumps({
"model": "gpt-4o-mini",
"messages": [
{"role": "system", "content": system_prompt},
{"role": "user", "content": user_prompt}
],
"temperature": 0.1
}).encode("utf-8")
req = urllib.request.Request(
"https://api.openai.com/v1/chat/completions",
data=payload,
headers={"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
)
with urllib.request.urlopen(req) as resp:
data = json.loads(resp.read().decode("utf-8"))
return data["choices"][0]["message"]["content"].strip()
def create_self_healing_pr(repo: str, branch: str, file_path: str, updated_content: str, token: str):
"""Applies remediation by opening a Git branch and PR via GitHub REST API."""
headers = {
"Authorization": f"Bearer {token}",
"Accept": "application/vnd.github.v3+json"
}
base_url = f"https://api.github.com/repos/{repo}"
# 1. Fetch default branch SHA
req = urllib.request.Request(f"{base_url}/git/ref/heads/main", headers=headers)
with urllib.request.urlopen(req) as resp:
sha = json.loads(resp.read().decode())["object"]["sha"]
# 2. Create target branch
branch_payload = json.dumps({"ref": f"refs/heads/{branch}", "sha": sha}).encode()
req = urllib.request.Request(f"{base_url}/git/refs", data=branch_payload, headers=headers)
urllib.request.urlopen(req)
# 3. Fetch file SHA
req = urllib.request.Request(f"{base_url}/contents/{file_path}?ref={branch}", headers=headers)
with urllib.request.urlopen(req) as resp:
file_sha = json.loads(resp.read().decode())["sha"]
# 4. Commit updated file
encoded_content = base64.b64encode(updated_content.encode()).decode()
commit_payload = json.dumps({
"message": "fix(ci): autonomous self-healing patch for test failure",
"content": encoded_content,
"sha": file_sha,
"branch": branch
}).encode()
req = urllib.request.Request(f"{base_url}/contents/{file_path}", data=commit_payload, headers=headers, method="PUT")
urllib.request.urlopen(req)
# 5. Open Pull Request
pr_payload = json.dumps({
"title": "🤖 [Self-Healing] Fix CI Failure Regression",
"head": branch,
"base": "main",
"body": "This Pull Request was automatically generated by the Autonomous CI Log Triage Engine after detecting deterministic failures in the pipeline."
}).encode()
req = urllib.request.Request(f"{base_url}/pulls", data=pr_payload, headers=headers)
urllib.request.urlopen(req)
4. Deployment, Timeouts, and Rate Limits
To run this in production without adding latency to green runs:
-
Conditional Triggers: Only execute the triage step on
if: failure()in GitHub Actions orwhen: on_failurein GitLab CI. - Run-Time Budgets: Restrict the DuckDB scan window to the last 5,000 lines. DuckDB handles this in under 200ms.
-
LLM Cost Guardrails: Strip ANSI escape codes (
\x1b\[[0-9;]*m) before sending payloads to LLMs to prevent token inflation. Truncate context to include only the stack trace boundary and the active function block. -
Prevent PR Flooding: Cache the hash of the error signature (
MD5(trace_text)) in your build cache. If an identical PR is already pending, skip recreation.
5. Conclusion & Ready-to-Use Workflow
You can implement this architecture using the code snippets above or integrate it into an automated n8n webhook listener. By replacing third-party SaaS log aggregators with embedded DuckDB queries, you gain deterministic triage while keeping logs completely inside your security perimeter.
If you prefer a plug-and-play solution, you can get the complete turnkey package containing production GitHub Action workflows, Docker container configurations, synthetic log test fixtures, and pre-built n8n triage workflows:
- Instant Access on Whop
- Direct Download on Gumroad (Use promo code EARLYBIRD for 20% off)
Top comments (0)