DEV Community

Site Reliability Engineering

Site Reliability Engineering principles, practices, and culture.

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
One Second Without DNS, Eight Hours Offline

One Second Without DNS, Eight Hours Offline

2
Comments
5 min read
Capacity Planning a SaaS Chatbot API Across OpenAI, Claude, and Gemini

Capacity Planning a SaaS Chatbot API Across OpenAI, Claude, and Gemini

Comments
8 min read
💀 The Most Dangerous Deletions Don't Break the Build. They Break the Deploy. An AI-Assisted Recovery Story 🤖

💀 The Most Dangerous Deletions Don't Break the Build. They Break the Deploy. An AI-Assisted Recovery Story 🤖

Comments
7 min read
One bad Kafka record shouldn't crash a Flink Stateful Functions job

One bad Kafka record shouldn't crash a Flink Stateful Functions job

Comments
3 min read
How to monitor third-party provider outages with two signals

How to monitor third-party provider outages with two signals

Comments
4 min read
AI Will Be Wrong Sometimes. What Then?

AI Will Be Wrong Sometimes. What Then?

Comments 1
3 min read
AWS & SRE Field Manual (Part 8): Amazon CloudWatch Architecture, Telemetry Pipelines & Production Observability

AWS & SRE Field Manual (Part 8): Amazon CloudWatch Architecture, Telemetry Pipelines & Production Observability

5
Comments 1
4 min read
Scale Enterprise RAG: Vector DBs & Hybrid Search on Bare Metal

Scale Enterprise RAG: Vector DBs & Hybrid Search on Bare Metal

1
Comments 2
4 min read
Building a Reliable Public Results Pipeline When Sources Update at Different Times

Building a Reliable Public Results Pipeline When Sources Update at Different Times

Comments
4 min read
EU Speech-to-Text API Procurement: Compare Startup Pricing, SLOs, and Exit Tests

EU Speech-to-Text API Procurement: Compare Startup Pricing, SLOs, and Exit Tests

Comments
7 min read
Our EKS Node Crashed and Four Auto-Recovery Mechanisms All Failed. Here's Why.

Our EKS Node Crashed and Four Auto-Recovery Mechanisms All Failed. Here's Why.

Comments
7 min read
I Built a Disaster Recovery Tool Because Row Counts Lied to Me

I Built a Disaster Recovery Tool Because Row Counts Lied to Me

2
Comments 1
5 min read
The DevOps interview question that predicts an outage

The DevOps interview question that predicts an outage

Comments
5 min read
Idempotency: The Bug You Don't Notice Until Production

Idempotency: The Bug You Don't Notice Until Production

Comments
7 min read
Kill switch for noisy uptime checks: a feature flag to disable a polling client

Kill switch for noisy uptime checks: a feature flag to disable a polling client

Comments
7 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.