DEV Community

Chris Lee
Chris Lee

Posted on

The 3am Lesson: Observability Isn't Optional, It's Infrastructure

I spent six hours last week chasing a latency spike that only appeared when our API hit 2,000 RPS. The metrics dashboard showed healthy CPU, memory, and response times—until it didn't. The culprit? A synchronous HTTP call to an internal service buried three layers deep in a middleware chain, with no timeout configured. Under load, thread pools exhausted silently. Requests queued. The service didn't crash; it just... stopped responding. We had logs, but they were unstructured text. We had metrics, but no traces. We had alerts, but they fired after users complained.

The hard truth: you cannot debug what you cannot see, and you cannot see what you didn't instrument before the fire started. "We'll add tracing later" is the same lie as "we'll write tests later." Distributed tracing, structured logging with correlation IDs, and SLO-based alerting aren't nice-to-haves—they're the load-bearing walls of a scalable system. Now, every new service gets OpenTelemetry baked into the scaffold. Every PR requires a trace_id in logs. We treat observability like schema migrations: mandatory, versioned, and tested in CI. The next 3am page will still hurt, but at least I'll know where to look.

Top comments (0)