DEV Community

Site Reliability Engineering

Site Reliability Engineering principles, practices, and culture.

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
SLA vs SLO vs SLI: what's the difference and why it matters

SLA vs SLO vs SLI: what's the difference and why it matters

Comments
9 min read
SLO examples for financial services: what good performance looks like in fintech

SLO examples for financial services: what good performance looks like in fintech

Comments
6 min read
OperatorMesh: Incident Triage Without Dashboard Noise

OperatorMesh: Incident Triage Without Dashboard Noise

Comments
1 min read
S3 Is Starting to Feel Like a File System — But Not Quite

S3 Is Starting to Feel Like a File System — But Not Quite

1
Comments
2 min read
CI/CD Auto-Remediation: The Complete Guide for SRE and Platform Teams (2026)

CI/CD Auto-Remediation: The Complete Guide for SRE and Platform Teams (2026)

2
Comments 1
12 min read
Closed-Loop SRE for Kubernetes: Auto-Remediating Pod Crashloops Before the On-Call Pages

Closed-Loop SRE for Kubernetes: Auto-Remediating Pod Crashloops Before the On-Call Pages

1
Comments
6 min read
My First dev.to Post — And a 1-Evening SRE System That Changed Our On-Call

My First dev.to Post — And a 1-Evening SRE System That Changed Our On-Call

Comments
2 min read
Your Kubernetes backups are lying to you

Your Kubernetes backups are lying to you

Comments
4 min read
subPath ConfigMap Mounts Don't Hot-Reload: Silent Drift in Kubernetes

subPath ConfigMap Mounts Don't Hot-Reload: Silent Drift in Kubernetes

Comments
6 min read
Human Operators in Distributed Financial Systems: When People Become Part of the Architecture

Human Operators in Distributed Financial Systems: When People Become Part of the Architecture

Comments
4 min read
80% of GitHub Repos Still Use Static AWS Credentials in 2026

80% of GitHub Repos Still Use Static AWS Credentials in 2026

Comments
4 min read
How to Fixed a Kubernetes CrashLoopBackOff in Production

How to Fixed a Kubernetes CrashLoopBackOff in Production

1
Comments
2 min read
Incident response / On-call: timeouts — operational runbook (playbook thực chiến)

Incident response / On-call: timeouts — operational runbook (playbook thực chiến)

Comments
3 min read
From MVP to Production: Scaling a Speech AI Service

From MVP to Production: Scaling a Speech AI Service

Comments
3 min read
I Don't Want AI to Replace DevOps. I Want It to Read the Docs I'm Too Tired to Read

I Don't Want AI to Replace DevOps. I Want It to Read the Docs I'm Too Tired to Read

4
Comments
9 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.