DEV Community

Site Reliability Engineering

Site Reliability Engineering principles, practices, and culture.

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Setting up a local Kubernetes cluster with minikube

Setting up a local Kubernetes cluster with minikube

1
Comments
14 min read
The SIGTERM our build workers ignored, and the 90s that fixed it

The SIGTERM our build workers ignored, and the 90s that fixed it

Comments
4 min read
Why Most Disaster Recovery Tests Don't Test Recovery

Why Most Disaster Recovery Tests Don't Test Recovery

Comments
4 min read
Claude Code detected a hack that never happened, then spiraled

Claude Code detected a hack that never happened, then spiraled

4
Comments 5
5 min read
Security Monitoring for SRE Teams

Security Monitoring for SRE Teams

Comments
2 min read
Provider drift broke our regression evals. We pinned versions through Bifrost.

Provider drift broke our regression evals. We pinned versions through Bifrost.

Comments
4 min read
Beyond DORA: A Five-Metric Framework for SRE Maturity in Regulated Enterprises

Beyond DORA: A Five-Metric Framework for SRE Maturity in Regulated Enterprises

Comments
13 min read
Docker Containerization Habits That Keep Production Calm

Docker Containerization Habits That Keep Production Calm

Comments
6 min read
Debugging Containers From the Terminal: A Practical Docker CLI Workflow

Debugging Containers From the Terminal: A Practical Docker CLI Workflow

1
Comments
7 min read
The 60-Second Break-Glass Protocol: Hot-Patching Live Production Outages via Local Tunnels

The 60-Second Break-Glass Protocol: Hot-Patching Live Production Outages via Local Tunnels

Comments
11 min read
Why Your Microservices Need Circuit Breakers (And How to Add Them)

Why Your Microservices Need Circuit Breakers (And How to Add Them)

Comments 2
2 min read
Instrumenting Legacy Code Without Rewriting It

Instrumenting Legacy Code Without Rewriting It

Comments
2 min read
System Design - Availability & Reliability: What "99.9% Uptime" Really Means (And Why It's Not Enough)

System Design - Availability & Reliability: What "99.9% Uptime" Really Means (And Why It's Not Enough)

Comments
6 min read
Configure Audit Logging in Kubernetes

Configure Audit Logging in Kubernetes

Comments
4 min read
The 54-point production deployment checklist that saves you from 3am rollbacks

The 54-point production deployment checklist that saves you from 3am rollbacks

Comments
3 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.