DEV Community

Site Reliability Engineering

Site Reliability Engineering principles, practices, and culture.

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
What I Actually Pay For When My LLM Bill Doubles Overnight

What I Actually Pay For When My LLM Bill Doubles Overnight

Comments
4 min read
Logging & Observability Best Practices from Bronto

Logging & Observability Best Practices from Bronto

2
Comments
6 min read
How I Built FRIDAY ? An Autonomous Incident Investigation Agent That Reduced MTTR by 65%

How I Built FRIDAY ? An Autonomous Incident Investigation Agent That Reduced MTTR by 65%

1
Comments 3
8 min read
What 99.9% vs 99.99% Uptime Really Means: An SRE Reality Check

What 99.9% vs 99.99% Uptime Really Means: An SRE Reality Check

Comments
3 min read
What Is Multi-Agent SRE? A Practical Introduction

What Is Multi-Agent SRE? A Practical Introduction

Comments
3 min read
Surviving an AZ Failover for Our Build Runner Fleet at 3am

Surviving an AZ Failover for Our Build Runner Fleet at 3am

Comments
4 min read
Why Fail-Closed Security Matters for Critical Systems

Why Fail-Closed Security Matters for Critical Systems

1
Comments
1 min read
The Cost Math Behind Our CI Cache Hit Rate Going From 40% to 91%

The Cost Math Behind Our CI Cache Hit Rate Going From 40% to 91%

Comments
4 min read
Railway vs AWS: When Leaving Railway Means Owning Reliability

Railway vs AWS: When Leaving Railway Means Owning Reliability

2
Comments 2
14 min read
The Future of SRE: What the Next 5 Years Look Like

The Future of SRE: What the Next 5 Years Look Like

Comments
3 min read
How we caught a silent IO storm before it hit production 🌩️

How we caught a silent IO storm before it hit production 🌩️

Comments
1 min read
Great Stack to Doesn't Work #10 — Season Finale: "When PagerDuty Calls at 3 AM"

Great Stack to Doesn't Work #10 — Season Finale: "When PagerDuty Calls at 3 AM"

Comments
12 min read
Chaos testing your CI runner fleet when half the jobs call an LLM

Chaos testing your CI runner fleet when half the jobs call an LLM

Comments 1
4 min read
Why Applications Work Locally But Fail in Production

Why Applications Work Locally But Fail in Production

Comments
4 min read
A Clean-Room Kubernetes CrashLoopBackOff Incident Exercise for SRE/DevOps Learners

A Clean-Room Kubernetes CrashLoopBackOff Incident Exercise for SRE/DevOps Learners

Comments
4 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.