DEV Community

Esa Data
Esa Data

Posted on

From Telemetry to Automated Incident Response: My DevOps & Observability Journey

Observability and Incident Response Pipeline

Description: A hands-on learning journey building observability and automated incident response into a FastAPI order-tracking application using OpenTelemetry, Prometheus, Loki, Tempo, Grafana, and Codex CLI.

From Telemetry to Automated Incident Response

As part of the DataTalksClub AI Dev Tools Zoomcamp, I worked on a small but practical challenge: turning a simple FastAPI application into an observable system with automated incident response.

The project started as a basic Order Tracker application backed by SQLite.

By the end, it had evolved into a system that could:

  • collect metrics, logs, and traces
  • visualize telemetry in Grafana
  • detect application errors through alerting
  • send incidents to an incident-response service
  • provide evidence for investigation
  • trigger a coding agent to investigate and remediate an issue

The most valuable lesson for me was that observability is not just about collecting data.

It is about connecting detection → investigation → remediation → verification.

The Problem

A basic application can work perfectly during development and still fail in production.

For an order-tracking API, a failure might look simple:

GET /api/orders/express-1002
Enter fullscreen mode Exit fullscreen mode

returns a server error.

Without observability, the workflow often becomes:

Something is broken
        ↓
Check application logs
        ↓
Search for the error
        ↓
Try to reproduce it
        ↓
Find the problematic code
        ↓
Apply a fix
        ↓
Restart the application
        ↓
Test again
Enter fullscreen mode Exit fullscreen mode

This works, but it becomes increasingly difficult as the system grows.

I wanted to build a workflow where telemetry could help connect these steps.

The Solution

The final architecture introduced several components:

Observability and Incident Response Pipeline

This architecture separates telemetry collection, storage, visualization, alerting, and incident response.

Instrumenting the Application

The first step was adding OpenTelemetry instrumentation to the FastAPI application.

I wanted each order lookup to produce useful telemetry, including:

  • HTTP route
  • HTTP status code
  • request information
  • logs
  • traces
  • request metrics

For example:

GET /api/orders/standard-1001
→ 200
Enter fullscreen mode Exit fullscreen mode

and:

GET /api/orders/standard-1002
→ 404
Enter fullscreen mode Exit fullscreen mode

The important part was not simply knowing that a request happened.

The telemetry needed enough context to answer:

What endpoint failed, what was the HTTP status, and what happened during the request?

Adding the Observability Stack

The next step was connecting the application to an observability stack:

  • OpenTelemetry Collector — receives and routes telemetry
  • Prometheus — metrics
  • Loki — logs
  • Tempo — traces
  • Grafana — visualization and alerting

The OpenTelemetry Collector became the central telemetry pipeline.

Conceptually:

Application
    │
    │ OTLP
    ▼
OpenTelemetry Collector
    │
    ├── Metrics → Prometheus
    │
    ├── Logs    → Loki
    │
    └── Traces  → Tempo
Enter fullscreen mode Exit fullscreen mode

Grafana then provides a single interface for exploring these signals.

This was one of the important lessons from the project:

Metrics tell you that something is happening.
Logs help explain what happened.
Traces help show where it happened.

Using the three together provides much more useful operational context than relying on a single signal.

Grafana Alerting

After telemetry was available, I configured Grafana alerting for server-side failures.

The goal was to detect 5xx responses.

A 404 response such as:

GET /api/orders/standard-1002
→ 404
Enter fullscreen mode Exit fullscreen mode

should not trigger the 5xx incident workflow.

This distinction was useful because it forced me to think about alert conditions rather than simply alerting on every failed request.

The alert also needed to provide enough information for the next stage of the workflow.

Automated Incident Response

The next step was adding a dedicated incident-response service.

The service exposes:

POST /alerts
Enter fullscreen mode Exit fullscreen mode

When Grafana sends an alert, the responder stores the incident information and gathers useful evidence.

The idea is to transform:

Grafana Alert
Enter fullscreen mode Exit fullscreen mode

into:

Incident
├── endpoint
├── status
├── logs
├── traces
└── investigation context
Enter fullscreen mode Exit fullscreen mode

This creates a bridge between observability and automated remediation.

Connecting an AI Coding Agent

The most interesting part of the exercise was connecting the incident workflow to a coding agent.

Instead of simply notifying a developer:

🚨 500 error detected
Enter fullscreen mode Exit fullscreen mode

the system can provide the coding agent with context about the incident.

The intended workflow becomes:

5xx detected
     ↓
Grafana alert
     ↓
Incident-response service
     ↓
Collect evidence
     ↓
Coding agent investigates
     ↓
Identify root cause
     ↓
Apply fix
     ↓
Restart application
     ↓
Verify behavior
Enter fullscreen mode Exit fullscreen mode

This changes the role of AI from simply generating code to participating in an operational workflow.

Finding the Bug

The final exercise intentionally exposed a bug in the Express order lookup.

The problematic implementation calculated the estimated delivery date using:

estimated_at = placed_at.replace(day=placed_at.day + 2)
Enter fullscreen mode Exit fullscreen mode

This looks reasonable at first glance.

However, it assumes that the resulting day exists in the same month.

For an order created near the end of a month, that assumption can fail.

The fix was to use date arithmetic instead:

estimated_at = placed_at + timedelta(days=2)
Enter fullscreen mode Exit fullscreen mode

This allows Python's datetime implementation to correctly handle month boundaries.

The important part was not only finding the line of code.

The observability and incident-response workflow provided the path from:

HTTP failure
    ↓
Telemetry
    ↓
Alert
    ↓
Incident evidence
    ↓
Root-cause investigation
    ↓
Code change
    ↓
Verification
Enter fullscreen mode Exit fullscreen mode

What I Learned

1. Observability is more than logging

Before this project, it was easy to think of observability as:

"Add some logs."
Enter fullscreen mode Exit fullscreen mode

The project changed that perspective.

A useful observability system combines multiple signals and makes them actionable.

2. Alerts need context

An alert saying:

Something went wrong
Enter fullscreen mode Exit fullscreen mode

is not particularly useful.

An actionable incident should provide enough context to start an investigation.

For example:

Endpoint
HTTP status
Logs
Trace
Time window
Dashboard context
Enter fullscreen mode Exit fullscreen mode

3. Automation should be connected to verification

Automatically changing code is not enough.

A remediation workflow should also verify that the application works after the change.

That makes the workflow closer to:

Detect → Investigate → Fix → Verify
Enter fullscreen mode Exit fullscreen mode

rather than simply:

Detect → Fix
Enter fullscreen mode Exit fullscreen mode

4. AI agents need operational context

A coding agent becomes much more useful when it receives evidence from the system instead of being asked to investigate blindly.

Telemetry can provide the context needed for the agent to understand what actually happened.

Technology Stack

The project brought several tools together:

Area Technology
API FastAPI
Runtime Python
Database SQLite
Telemetry OpenTelemetry
Telemetry Pipeline OpenTelemetry Collector
Metrics Prometheus
Logs Loki
Traces Tempo
Visualization Grafana
Containers Docker / Docker Compose
AI-assisted remediation Codex CLI

Verification

The completed workflow was validated through the homework tasks.

Step Scenario Result
Q1 Application health check 200
Q2 Existing order lookup 200
Q3 Missing order lookup 404
Q4 5xx alert behavior Normal for 404
Q5 Incident-response workflow Completed
Q6 Express order incident Root cause identified and fixed

The important outcome was not just that each individual component worked.

The complete workflow worked as a chain.

Final Takeaway

The biggest lesson from this project was that observability becomes much more valuable when it is connected to action.

A mature workflow can look like:

Application
     ↓
Telemetry
     ↓
Observability
     ↓
Alerting
     ↓
Incident Response
     ↓
AI-assisted Investigation
     ↓
Remediation
     ↓
Verification
Enter fullscreen mode Exit fullscreen mode

This project gave me hands-on experience connecting these pieces together rather than learning them as isolated technologies.

For me, that's the real value of Learning in Public: documenting not only what I built, but also the problems I encountered, the root causes I found, and what changed in my understanding along the way.


Acknowledgments

This project was completed as part of the DataTalksClub AI Dev Tools Zoomcamp — Homework 4: DevOps and Observability for AI-Built Apps.

Thanks to the DataTalksClub community for creating practical exercises that connect software development, DevOps, observability, and AI-assisted engineering.


Github Repository: order-tracker-v2

Top comments (0)