DEV Community

Site Reliability Engineering

Site Reliability Engineering principles, practices, and culture.

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
The On-Call Schedule Math Nobody Does

The On-Call Schedule Math Nobody Does

Comments 1
2 min read
Putting an LLM Gateway in Front of Our Build Agents: Why We Picked Bifrost

Putting an LLM Gateway in Front of Our Build Agents: Why We Picked Bifrost

Comments
4 min read
Platform Engineering in Practice: Hardening Backstage with SRE Watchdogs, Zero-Touch RBAC, and SSRF Mitigation

Platform Engineering in Practice: Hardening Backstage with SRE Watchdogs, Zero-Touch RBAC, and SSRF Mitigation

Comments
4 min read
Your OTel Traces Are Lying to You Observability for the Reasoning Layer

Your OTel Traces Are Lying to You Observability for the Reasoning Layer

Comments 1
5 min read
Production Log Parsing Patterns That Break Real Kubernetes Clusters (and How to Fix Them)

Production Log Parsing Patterns That Break Real Kubernetes Clusters (and How to Fix Them)

Comments
4 min read
The Future Guide for Escaping Single-Provider Administrative Failure

The Future Guide for Escaping Single-Provider Administrative Failure

Comments
6 min read
I Made 4 LLMs Argue With Each Other to Write Better Runbooks. Here's What Happened.

I Made 4 LLMs Argue With Each Other to Write Better Runbooks. Here's What Happened.

Comments
5 min read
Chaos Engineering Is Theater Without These Three Things

Chaos Engineering Is Theater Without These Three Things

1
Comments
2 min read
We're hiring a DevOps Content Engineer – Remote LATAM

We're hiring a DevOps Content Engineer – Remote LATAM

2
Comments
1 min read
The Runbook Is Already Lying to you.

The Runbook Is Already Lying to you.

Comments
8 min read
The Architecture — Prometheus, Grafana, and StatsD for Batch Workloads

The Architecture — Prometheus, Grafana, and StatsD for Batch Workloads

Comments
5 min read
Debugging Production Alerts Without Chasing The Wrong Problem

Debugging Production Alerts Without Chasing The Wrong Problem

Comments
2 min read
Why Developers Should Learn How Systems Fail

Why Developers Should Learn How Systems Fail

Comments
3 min read
How I Taught My Incident Alerts to Say "This Broke 3 Minutes After Your Last Deploy"

How I Taught My Incident Alerts to Say "This Broke 3 Minutes After Your Last Deploy"

Comments 2
3 min read
Choosing Your First SLI: A Decision Framework for New SRE Teams

Choosing Your First SLI: A Decision Framework for New SRE Teams

Comments
2 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.