DEV Community

Ayi NEDJIMI
Ayi NEDJIMI

Posted on

How to Build an AI-Powered Log Summarizer for DevOps

Debugging a production incident means combing through hundreds of thousands of log lines. By the time you find the root cause, the outage has already cost you an hour. An LLM-backed log summarizer doesn't replace your observability stack — but it gives your on-call engineer a head start that matters.

The Problem with Raw Logs

Log aggregation tools like Loki, Elasticsearch, and Datadog are excellent at storing and searching logs, but they leave interpretation entirely to you. You run a query, you get 10,000 lines, and you start scrolling.

The gap isn't storage or retrieval — it's comprehension. You need someone (or something) to read a window of logs and say: "Between 14:23 and 14:31, the payment service returned 503 errors on every POST to /checkout. Three upstream timeout messages from the Redis connection pool preceded each failure."

A language model can produce exactly that — if you feed it the right data and give it the right prompt.

Architecture: What We're Building

The tool has three parts:

  1. A log collector — reads from a file, stdin, or a log query API
  2. A preprocessor — deduplicates noise, truncates to stay within token limits
  3. An LLM summarizer — sends the cleaned logs and returns a structured summary

We'll keep it as a Python CLI tool with no framework dependencies beyond httpx for the API call. One script, one responsibility.

Collecting and Preprocessing Logs

Raw logs have two problems: volume and repetition. Before sending anything to a language model, you need to reduce both.

import re
from collections import Counter
from typing import Iterator

MAX_LINES = 200
NOISE_PATTERNS = [
    re.compile(r'health_?check', re.IGNORECASE),
    re.compile(r'GET /ping'),
    re.compile(r'keepalive'),
]

def is_noise(line: str) -> bool:
    return any(p.search(line) for p in NOISE_PATTERNS)

def deduplicate(lines: list[str], threshold: int = 5) -> list[str]:
    """Keep first N occurrences of lines that repeat more than threshold times."""
    counts: Counter = Counter()
    result = []
    for line in lines:
        key = re.sub(r'\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}[^ ]*', '<ts>', line)
        key = re.sub(r'\b\d+\b', '<n>', key)
        counts[key] += 1
        if counts[key] <= threshold:
            result.append(line)
    return result

def preprocess(lines: Iterator[str]) -> str:
    cleaned = [l.rstrip() for l in lines if l.strip() and not is_noise(l)]
    deduped = deduplicate(cleaned)
    truncated = deduped[-MAX_LINES:]  # keep most recent
    return "\n".join(truncated)
Enter fullscreen mode Exit fullscreen mode

The deduplicate function normalizes timestamps and numbers before counting, so "connection refused at 14:23:01" and "connection refused at 14:23:07" count as the same pattern. You keep the first five occurrences, then stop — the language model doesn't need to see the same error 400 times to understand it's a pattern.

The MAX_LINES = 200 cap is conservative. With typical log verbosity, 200 lines fits in 3–4k tokens, leaving plenty of room for the model's response within a 16k context window.

Summarizing with an LLM

The summarization call is straightforward. The key is the prompt: you want the model to identify errors, correlate events, and surface the probable root cause — not paraphrase the logs.

import httpx
import os

API_URL = "https://api.openai.com/v1/chat/completions"  # any OpenAI-compatible endpoint

SYSTEM_PROMPT = """You are a DevOps incident assistant. Given a block of application logs,
produce a structured summary with these sections:
- Timeline: key events with approximate timestamps
- Errors: list of distinct error types and their count
- Root cause hypothesis: your best guess at what caused the issue
- Recommended next steps: 2-3 concrete investigation actions

Be concise. If you are uncertain, say so. Do not invent details not present in the logs."""

def summarize_logs(log_text: str, model: str = "gpt-4o-mini") -> str:
    token = os.environ["OPENAI_API_KEY"]
    payload = {
        "model": model,
        "messages": [
            {"role": "system", "content": SYSTEM_PROMPT},
            {"role": "user", "content": f"<logs>\n{log_text}\n</logs>"},
        ],
        "temperature": 0.2,
        "max_tokens": 800,
    }
    with httpx.Client(timeout=30) as client:
        resp = client.post(
            API_URL,
            headers={"Authorization": f"Bearer {token}"},
            json=payload,
        )
        resp.raise_for_status()
        return resp.json()["choices"][0]["message"]["content"]
Enter fullscreen mode Exit fullscreen mode

A few things worth noting:

  • temperature=0.2 keeps output factual. Higher values produce more varied but less reliable summaries.
  • The <logs> XML-style wrapper helps the model separate log content from instructions.
  • max_tokens=800 is enough for a useful summary without wasting budget on padding.

This function works with any OpenAI-compatible API, including local models served via Ollama or vLLM. Swap the API_URL and remove the auth header for local inference.

Putting It Together: A CLI Tool

#!/usr/bin/env python3
import sys
import argparse
from pathlib import Path

def main():
    parser = argparse.ArgumentParser(description="Summarize logs with an LLM")
    parser.add_argument("file", nargs="?", help="Log file (default: stdin)")
    parser.add_argument("--model", default="gpt-4o-mini")
    args = parser.parse_args()

    if args.file:
        lines = Path(args.file).read_text().splitlines(keepends=True)
    else:
        lines = sys.stdin

    log_text = preprocess(lines)
    if not log_text.strip():
        print("No meaningful log content after preprocessing.", file=sys.stderr)
        sys.exit(1)

    summary = summarize_logs(log_text, model=args.model)
    print(summary)

if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

Run it against a file or pipe a live stream directly:

# From a file
python log_summarizer.py /var/log/app/production.log

# From a live Kubernetes tail
kubectl logs -n production deploy/api-server --since=10m | python log_summarizer.py
Enter fullscreen mode Exit fullscreen mode

The kubectl pipe pattern is where this shines. During an active incident, you pipe the last 10 minutes of logs and get a structured summary in about 3 seconds — faster than reading the same window yourself, with output you can paste directly into your incident channel.

Practical Limits to Know Before You Ship

Log sensitivity. Application logs frequently contain PII, session tokens, or internal IP addresses. Before sending logs to a hosted language model, run a redaction pass. A basic regex sweep for email patterns, bearer tokens, and internal hostnames is enough for most cases. For a checklist covering what to scrub before sending logs to any external service, the security hardening checklists at AYI NEDJIMI Consultants include a data-before-LLM section.

Context window exhaustion. The MAX_LINES cap is a safety rail, but measure token count before sending if you're dealing with verbose services. Libraries like tiktoken (for OpenAI models) let you count tokens before the API call and bail early if the window would overflow.

Hallucination on sparse data. If your preprocessed log window is only 15 lines, the model has almost nothing to work with and will often produce a plausible-sounding but fabricated root cause. Add a minimum threshold: if fewer than 20 lines survive preprocessing, output them raw and skip the LLM call entirely.

Cost. With gpt-4o-mini at roughly $0.15 per 1M input tokens, 200 lines of logs (about 4k tokens) costs less than a tenth of a cent per invocation. Running this on every deploy or on-call page is fine. Running it in a hot loop on every HTTP error is not.

The Takeaway

The hard part isn't the code — it's getting the system prompt to reliably produce structured output. Start with a fixed log source, a single service, and iterate on the prompt until the model consistently delivers the timeline/errors/hypothesis format you want. Then generalize.

This also works as an onboarding tool: new engineers can pipe an unfamiliar service's logs and get a plain-language explanation of what's happening without needing to know every log format on day one. The 80 lines of Python takes an afternoon; the prompt takes longer — but once it's tuned, it earns its keep in the first real incident.


I run AYI NEDJIMI Consultants, a cybersecurity consulting firm. We publish free security hardening checklists — PDF and Excel.

Top comments (0)