DEV Community

Richard Atodo
Richard Atodo

Posted on

๐Ÿ› ๏ธ Building Multi-Source Log Ingestion for an AI DevOps Incident Copilot

๐Ÿ“Œ Introduction

One of the first challenges in building an incident investigation tool is getting logs from different systems into a form the application can work with consistently.

  • Nginx โ†’ access and error logs
  • Kubernetes โ†’ events about workloads and containers
  • Docker โ†’ container output
  • Applications โ†’ structured/unstructured messages
  • GitHub Actions โ†’ workflow and job failures

These sources differ in structure and terminology, but an incident investigation system needs to examine them together.

Milestone 3 of IncidentCopilot focused on building the ingestion foundation โ€” ensuring the application can accept, validate, normalize, and persist logs reliably.


๐Ÿ”Ž The Scope: Five Log Sources

Supported sources:

  • Nginx
  • Kubernetes
  • Docker
  • Application logs
  • GitHub Actions

Each source has its own parser module. Parsers produce a shared normalized representation with fields like:

  • Source
  • Timestamp
  • Severity
  • Service
  • Event type
  • Message
  • Metadata
  • Raw message

Source-specific details remain in metadata (e.g., Nginx โ†’ HTTP status code, Kubernetes โ†’ namespace + pod info).


โš™๏ธ Ingestion API

Endpoints:

Method Endpoint Purpose
POST /api/v1/logs Submit one log
POST /api/v1/logs/batch Submit multiple logs
GET /api/v1/logs List logs with pagination
GET /api/v1/logs/{id} Retrieve a specific log
  • Batch endpoint accepts 1โ€“100 records.
  • Source-specific schemas validate incoming data.
  • Invalid batch items โ†’ request rejected before ingestion.

โœ… What Worked โ€” and What Needed Fixing

  • Live API checks confirmed single + batch ingestion worked (201 Created, 200 OK).
  • But one test failed: expected empty collection โ†’ got six records.
  • Cause: test queried dev DB with leftover data from live checks.
  • Fix: test environment isolation.

๐Ÿงช Isolating the Test Database

Solution:

  • Created dedicated DB โ†’ incidentcopilot_test.
  • Applied Alembic migrations.
  • Added pytest fixture in backend/tests/conftest.py:
    • Test-specific SQLAlchemy engine + session
    • Overrides FastAPI DB dependency
    • Transaction rollback after each test

Benefits:

  1. Environment isolation โ†’ tests no longer run against dev DB.
  2. Test isolation โ†’ records rolled back after each test.

๐Ÿ” Verification

After correction:

Verification Result
First run 27 passed in 0.88s
Second run 27 passed in 2.00s
Alembic check No new upgrade ops
Git diff check No whitespace errors

Suite covers API, DB connectivity, health checks, ingestion validation, batch handling, unknown IDs.


๐Ÿšซ What This Milestone Does Not Do

Milestone 3 = ingestion foundation.

It does not yet:

  • Correlate logs into incidents
  • Build investigation timelines
  • Query vector DB
  • Generate AI diagnoses

๐Ÿ“š Lessons Learned

  1. Different sources need different validation, but a common representation.
  2. Live API checks โ‰  automated tests. Controlled setup matters.
  3. Test isolation improves engineering quality. Dedicated test DB + rollback fixtures made the suite reliable.

๐Ÿ”ฎ Whatโ€™s Next?

Next milestone โ†’ log normalization + correlation.

That work will connect related events into incident context. AI diagnosis and retrieval layers will come later.


๐Ÿ Conclusion

Milestone 3 delivered:

  • Multi-source ingestion workflow
  • Source-specific parsers + validation
  • Normalization with raw evidence preserved
  • Single + batch ingestion endpoints
  • PostgreSQL persistence
  • Reliable test isolation (27 passing tests)

This milestone wasnโ€™t just about accepting logs โ€” it was about making the foundation testable enough to trust as the project grows.


๐Ÿ”— GitHub: github.com/richardatodo/incidentcopilot

โžก๏ธ Next: Milestone 4 โ€” Normalization & Correlation Engine

Top comments (0)