DEV Community

Site Reliability Engineering

Site Reliability Engineering principles, practices, and culture.

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
When one reliability surface has to satisfy everyone

When one reliability surface has to satisfy everyone

1
Comments
5 min read
DevOps vs SRE: Key Differences Explained [2026 Guide]

DevOps vs SRE: Key Differences Explained [2026 Guide]

Comments
2 min read
Scaling On-Call When You Only Have 5 Engineers

Scaling On-Call When You Only Have 5 Engineers

Comments
2 min read
Humanizing Artificial Intelligence for SRE Teams: Reducing Alert Fatigue With Smarter AI Guidance

Humanizing Artificial Intelligence for SRE Teams: Reducing Alert Fatigue With Smarter AI Guidance

2
Comments 2
9 min read
A note on building reliability infrastructure for AI agents and why post-incident debugging matters more than pre-flight validation.

A note on building reliability infrastructure for AI agents and why post-incident debugging matters more than pre-flight validation.

1
Comments
4 min read
Why Most “Senior” DevOps Engineers Are Building Fragile Infrastructure — And Why the Industry Rewards It

Why Most “Senior” DevOps Engineers Are Building Fragile Infrastructure — And Why the Industry Rewards It

Comments
4 min read
Observability in 2026: Distributed Tracing Replaced Logs, and OpenTelemetry Won

Observability in 2026: Distributed Tracing Replaced Logs, and OpenTelemetry Won

Comments
5 min read
The Human-in-the-Loop SRE: Designing Automation Escalation Policies for AI-Assisted Operations

The Human-in-the-Loop SRE: Designing Automation Escalation Policies for AI-Assisted Operations

Comments 1
15 min read
Kubernetes Upgrades Without Downtime

Kubernetes Upgrades Without Downtime

Comments
2 min read
Root Cause Analysis Across Every Signal, On One Screen

Root Cause Analysis Across Every Signal, On One Screen

Comments
8 min read
Auto-verifying your AI-SRE's fixes (Part II): HolmesGPT end-to-end on a real cluster

Auto-verifying your AI-SRE's fixes (Part II): HolmesGPT end-to-end on a real cluster

20
Comments 3
5 min read
AI SRE and AI DevOps: different problems, one reliability stack

AI SRE and AI DevOps: different problems, one reliability stack

2
Comments
6 min read
The Infrastructure Team Is the Real Single Point of Failure

The Infrastructure Team Is the Real Single Point of Failure

Comments
7 min read
Reading the Prompt You Did Not Send: Detection at the Inference Boundary

Reading the Prompt You Did Not Send: Detection at the Inference Boundary

Comments
5 min read
Stop paying for idle GPUs in your CI: batching LLM eval jobs

Stop paying for idle GPUs in your CI: batching LLM eval jobs

Comments
4 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.