DEV Community

Site Reliability Engineering

Site Reliability Engineering principles, practices, and culture.

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
The Art of Writing a Good Post-Mortem

The Art of Writing a Good Post-Mortem

Comments
1 min read
Why We Stopped Using Log Aggregation for Everything

Why We Stopped Using Log Aggregation for Everything

Comments
1 min read
Design Around the Point You Cannot Undo

Design Around the Point You Cannot Undo

Comments
4 min read
What is a Forward Deployed Engineer? 8 months in the role.

What is a Forward Deployed Engineer? 8 months in the role.

Comments
5 min read
How We Reduced Our Deployment Failure Rate to Under 2%

How We Reduced Our Deployment Failure Rate to Under 2%

Comments
1 min read
Why Uptime Percentages Hide More Than They Reveal

Why Uptime Percentages Hide More Than They Reveal

Comments
4 min read
The Hidden Cost of Flaky Tests

The Hidden Cost of Flaky Tests

Comments
1 min read
Running Postgres at Scale: Lessons Learned

Running Postgres at Scale: Lessons Learned

1
Comments 1
2 min read
Observability for Serverless: What's Different

Observability for Serverless: What's Different

Comments
2 min read
Stop Retry Storms: Backoff Is Not Enough Without a Budget

Stop Retry Storms: Backoff Is Not Enough Without a Budget

Comments
6 min read
From DevOps to SRE: Making the Transition

From DevOps to SRE: Making the Transition

Comments
2 min read
The SRE Interview: Questions I Actually Ask

The SRE Interview: Questions I Actually Ask

Comments
1 min read
Incident Retrospectives Without Blame

Incident Retrospectives Without Blame

Comments
1 min read
Alert Fatigue: The Silent Productivity Killer

Alert Fatigue: The Silent Productivity Killer

Comments
1 min read
The Real Cost of Downtime for Startups (and How to Quantify It)

The Real Cost of Downtime for Startups (and How to Quantify It)

Comments
4 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.