DEV Community

Site Reliability Engineering

Site Reliability Engineering principles, practices, and culture.

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
EBS gp2 burst credits ran dry and our builds slowed to a crawl

EBS gp2 burst credits ran dry and our builds slowed to a crawl

1
Comments
4 min read
Amazon Linux 2 is EOL on June 30, 2026 — here's everything that breaks

Amazon Linux 2 is EOL on June 30, 2026 — here's everything that breaks

Comments
2 min read
Semantic caching our flaky-test summariser: 58% fewer LLM calls

Semantic caching our flaky-test summariser: 58% fewer LLM calls

Comments
4 min read
Automate creation of Amazon CloudWatch alarms

Automate creation of Amazon CloudWatch alarms

1
Comments
4 min read
relysam v2.0.0 a core Reliability Engineering platform with AI/ML enhancements

relysam v2.0.0 a core Reliability Engineering platform with AI/ML enhancements

Comments
2 min read
Chaos Engineering Is Theater Without These Three Things

Chaos Engineering Is Theater Without These Three Things

1
Comments
2 min read
Why "is it fully fixed?" has no honest answer — a small-server story

Why "is it fully fixed?" has no honest answer — a small-server story

Comments
2 min read
"이제 완전히 고쳐졌나?"에 정직한 답은 없다 — 작은 서버 이야기

"이제 완전히 고쳐졌나?"에 정직한 답은 없다 — 작은 서버 이야기

Comments
1 min read
Chaos Engineering for Node.js Without the Infrastructure

Chaos Engineering for Node.js Without the Infrastructure

Comments
4 min read
Choosing Your First SLI: A Decision Framework for New SRE Teams

Choosing Your First SLI: A Decision Framework for New SRE Teams

Comments
2 min read
How DNS Actually Works: Resolution Hierarchy, Caching, and Production Failure Modes

How DNS Actually Works: Resolution Hierarchy, Caching, and Production Failure Modes

Comments
15 min read
Runbook Hygiene: Why Yours Are Lying to You

Runbook Hygiene: Why Yours Are Lying to You

Comments
2 min read
What's the Most Annoying Part of Incident Response? I Built 5 AI Tools Trying to Solve It

What's the Most Annoying Part of Incident Response? I Built 5 AI Tools Trying to Solve It

Comments
1 min read
An 8-minute outage from a dead NLB and a JVM that cached DNS forever

An 8-minute outage from a dead NLB and a JVM that cached DNS forever

1
Comments
4 min read
Fault-injecting our LLM provider to trust Bifrost fallbacks

Fault-injecting our LLM provider to trust Bifrost fallbacks

Comments 1
4 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.