DEV Community

Site Reliability Engineering

Site Reliability Engineering principles, practices, and culture.

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
The status page you can't fake: measured uptime, not published

The status page you can't fake: measured uptime, not published

1
Comments
4 min read
Designing Actionable Alerting Systems to Avoid IT Alert Fatigue

Designing Actionable Alerting Systems to Avoid IT Alert Fatigue

Comments
4 min read
Read-Only First: A Safer Adoption Model for Agentic Platform Engineering

Read-Only First: A Safer Adoption Model for Agentic Platform Engineering

1
Comments
3 min read
From Tool Discovery to Real Execution: Verifying a Multi-Cluster MCP Path

From Tool Discovery to Real Execution: Verifying a Multi-Cluster MCP Path

1
Comments
3 min read
Designing ChatOps Sessions for Kubernetes Agents

Designing ChatOps Sessions for Kubernetes Agents

1
Comments
3 min read
Guardrails Before Write Access: Building Agentic Kubernetes Operations with Human Approval

Guardrails Before Write Access: Building Agentic Kubernetes Operations with Human Approval

1
Comments
6 min read
Multi-Region Failover: Lessons from Running It Hot

Multi-Region Failover: Lessons from Running It Hot

Comments
3 min read
When the Agent Is Wrong: Surface Real Kubernetes Errors, Not Model Guesses

When the Agent Is Wrong: Surface Real Kubernetes Errors, Not Model Guesses

1
Comments
3 min read
Disaster Recovery Drills That Actually Work

Disaster Recovery Drills That Actually Work

Comments
3 min read
Why Kubernetes Exec Is Secret-Adjacent in Agentic Operations

Why Kubernetes Exec Is Secret-Adjacent in Agentic Operations

2
Comments
3 min read
The SRE Talent Gap: Why the US Needs 10x More Reliability Engineers and How to Train Them

The SRE Talent Gap: Why the US Needs 10x More Reliability Engineers and How to Train Them

Comments
16 min read
Website Uptime Monitoring: From a 20-Line Bash Script to Production-Ready Alerts

Website Uptime Monitoring: From a 20-Line Bash Script to Production-Ready Alerts

Comments
4 min read
70 false alerts a day — why uptime monitors cry wolf — and what the fix costs

70 false alerts a day — why uptime monitors cry wolf — and what the fix costs

Comments
6 min read
Green CI working software

Green CI working software

1
Comments
3 min read
AWS puts gray zone failures into the EKS control loop

AWS puts gray zone failures into the EKS control loop

Comments
3 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.