DEV Community

Site Reliability Engineering

Site Reliability Engineering principles, practices, and culture.

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Read-Only First: A Safer Adoption Model for Agentic Platform Engineering

Read-Only First: A Safer Adoption Model for Agentic Platform Engineering

1
Comments 1
3 min read
Decoding System Observability: Building Transparent and Resilient Architectures

Decoding System Observability: Building Transparent and Resilient Architectures

Comments
2 min read
Multi-Region Failover: Lessons from Running It Hot

Multi-Region Failover: Lessons from Running It Hot

Comments
3 min read
Guardrails Before Write Access: Building Agentic Kubernetes Operations with Human Approval

Guardrails Before Write Access: Building Agentic Kubernetes Operations with Human Approval

1
Comments 1
6 min read
Fix Docker Exit Code 137 (OOMKilled): Why It Happens and How to Stop It

Fix Docker Exit Code 137 (OOMKilled): Why It Happens and How to Stop It

3
Comments
6 min read
Disaster Recovery Drills That Actually Work

Disaster Recovery Drills That Actually Work

Comments
3 min read
Website Uptime Monitoring: From a 20-Line Bash Script to Production-Ready Alerts

Website Uptime Monitoring: From a 20-Line Bash Script to Production-Ready Alerts

2
Comments
4 min read
We open-sourced the SRE judgment that doesn't fit in a system prompt

We open-sourced the SRE judgment that doesn't fit in a system prompt

Comments
3 min read
Google Published Their AI SRE Blueprint. Here's the Line-by-Line Mapping to What the Community Has Been Building

Google Published Their AI SRE Blueprint. Here's the Line-by-Line Mapping to What the Community Has Been Building

Comments
3 min read
Beyond Ingress: Why the Kubernetes Gateway API is the Future of Cloud Native Networking

Beyond Ingress: Why the Kubernetes Gateway API is the Future of Cloud Native Networking

1
Comments
6 min read
What 60+ Claude Code memory entries taught me about solo ops

What 60+ Claude Code memory entries taught me about solo ops

4
Comments 7
5 min read
Operational Debt in Distributed Financial Systems: When Temporary Workarounds Become Architecture

Operational Debt in Distributed Financial Systems: When Temporary Workarounds Become Architecture

Comments
7 min read
Error budgets when downtime costs money: reliability engineering for payment-critical systems

Error budgets when downtime costs money: reliability engineering for payment-critical systems

Comments
10 min read
Safe Operating Throughput (SOT) as a First-Class SRE Metric: Derivation and Operationalization

Safe Operating Throughput (SOT) as a First-Class SRE Metric: Derivation and Operationalization

Comments
17 min read
Why Your AKS Pods Keep Getting OOMKilled Even When CPU Looks Fine

Why Your AKS Pods Keep Getting OOMKilled Even When CPU Looks Fine

Comments
4 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.