๐ Introduction
When building an AI-powered DevOps incident investigation system, itโs tempting to start with the flashy features: log analysis, AI-generated root-cause explanations, and automated troubleshooting.
But those depend on something more fundamental: a reliable backend, a structured data model, and persistent storage.
As part of my IncidentCopilot project, Milestone 2 focused on establishing that foundation using FastAPI, PostgreSQL, SQLAlchemy, and Alembic.
๐ Project Overview
IncidentCopilot is an open-source project designed to help developers and DevOps engineers investigate incidents more efficiently.
Planned features include:
- Log ingestion
- Incident correlation
- Investigation timelines
- Evidence retrieval
- AI-assisted diagnosis
Principle: deterministic logic and reliable storage first, AI later.
1๏ธโฃ Integrating PostgreSQL
- Added PostgreSQL 16 to Docker Compose.
- Configured persistent storage with a named volume.
- Health checks via
pg_isready. - Backend waits for DB readiness before starting.
โ ๏ธ Key detail:
- Host dev โ
localhost - Container dev โ
postgres(Compose service name)
2๏ธโฃ Setting up SQLAlchemy
- Introduced SQLAlchemy for DB access.
- Shared engine + session factory.
-
get_dbdependency ensures sessions close after requests.
3๏ธโฃ Creating Domain Models
Six initial models:
- Incident โ title, severity, status, timestamps
- Log โ source, timestamp, severity, service, event type, message, metadata (JSONB)
- IncidentLog โ connects incidents โ logs, includes relevance scoring
- Diagnosis โ summary, root cause, confidence, category, raw response (JSONB)
- Evidence โ links diagnosis โ logs, with reasoning
- InvestigationStep โ description, status, position
4๏ธโฃ Managing Schema with Alembic
- Added Alembic for migrations.
- Initial migration created tables + relationships.
- Verified schema alignment with:
alembic check
Result:
No new upgrade operations detected.
5๏ธโฃ Adding API Endpoints
| Endpoint | Purpose |
|---|---|
GET /health |
Confirms app health |
GET /ready |
Checks DB connectivity |
GET /api/v1/incidents |
Lists incidents |
GET /api/v1/logs |
Lists logs |
Example readiness response:
{
"status": "ready"
}
6๏ธโฃ Testing the Foundation
- Tests for health, readiness, DB connectivity, model registration, empty endpoints.
- Result:
6 passed, 1 warning in 2.33s
- Verified Docker Compose config + DB migration state.
- Backend resolved
postgreshostname correctly.
7๏ธโฃ Lessons Learned
- Clear domain model matters โ incidents, logs, evidence, diagnoses, steps each have distinct roles.
- Keep config outside code โ easier portability.
- Use migrations early โ reproducible schema changes.
- Health โ readiness โ DB connectivity matters.
- Build incrementally โ donโt rush correlation or AI.
๐ฎ Whatโs Next?
Milestone 3 โ multi-source log ingestion.
Future milestones:
- Normalization
- Incident correlation
- Investigation timelines
- Retrieval-augmented generation (RAG)
- AI-assisted diagnosis
๐ Conclusion
Milestone 2 wasnโt about flashy AI. It was about laying a dependable foundation.
By integrating PostgreSQL, SQLAlchemy, Alembic, structured domain models, versioned API endpoints, and automated tests, IncidentCopilot now has the backend infrastructure to grow.
Next challenge: make the foundation useful by ingesting operational logs.
๐ GitHub: github.com/richardatodo/incidentcopilot
โก๏ธ Next: Milestone 3 โ Multi-source Log Ingestion
Top comments (0)