DEV Community

Site Reliability Engineering

Site Reliability Engineering principles, practices, and culture.

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Why SLIs Matter More Than SLOs

Why SLIs Matter More Than SLOs

1
Comments
1 min read
SRE Body of Knowledge: A Practitioner's Annotated Reading List

SRE Body of Knowledge: A Practitioner's Annotated Reading List

Comments
17 min read
SRE Practices in Healthcare: Applying SLOs and Error Budgets to Life-Critical Systems

SRE Practices in Healthcare: Applying SLOs and Error Budgets to Life-Critical Systems

Comments
13 min read
How I Built a Chaos Engineering Platform for a Multi-Region Serverless Payment Aggregator (AWS + Quarkus + DDD)

How I Built a Chaos Engineering Platform for a Multi-Region Serverless Payment Aggregator (AWS + Quarkus + DDD)

Comments
5 min read
The PagerDuty Migration Playbook

The PagerDuty Migration Playbook

Comments
1 min read
Alarmas que despiertan por causa real, no por un nĂşmero random

Alarmas que despiertan por causa real, no por un nĂşmero random

Comments
9 min read
How We Cut Datadog Bills by 60% Without Losing Observability

How We Cut Datadog Bills by 60% Without Losing Observability

Comments
1 min read
The Archive Multiplier: Why eth_call at a Historical Block

The Archive Multiplier: Why eth_call at a Historical Block

Comments
9 min read
Building Your First Runbook: A Template That Actually Works

Building Your First Runbook: A Template That Actually Works

Comments
1 min read
Operational Consistency Is the Real Moat

Operational Consistency Is the Real Moat

Comments
3 min read
AIOps vs Traditional Monitoring: What Actually Changed

AIOps vs Traditional Monitoring: What Actually Changed

Comments
1 min read
Designing an On-Call Schedule That Doesn't Burn Out Your Team

Designing an On-Call Schedule That Doesn't Burn Out Your Team

1
Comments
3 min read
Building an Incident Response Playbook Library

Building an Incident Response Playbook Library

Comments
4 min read
Kubernetes Network Policies: Lessons from Production Incidents

Kubernetes Network Policies: Lessons from Production Incidents

Comments
4 min read
how to run a blameless postmortem that actually changes anything

how to run a blameless postmortem that actually changes anything

Comments
3 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.