DEV Community

Site Reliability Engineering

Site Reliability Engineering principles, practices, and culture.

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Observability Telemetry and Predictive AIOps

Observability Telemetry and Predictive AIOps

Comments
8 min read
MTTR Optimization: The 7 Levers That Actually Move the Needle

MTTR Optimization: The 7 Levers That Actually Move the Needle

Comments
3 min read
Service Maps: The Architectural Clarity Your Team Is Missing

Service Maps: The Architectural Clarity Your Team Is Missing

Comments
2 min read
AI in Incident Response: Hype vs. Reality in 2024

AI in Incident Response: Hype vs. Reality in 2024

Comments
3 min read
Kubernetes resource requests and limits explained: scheduling, throttling, and OOMKill

Kubernetes resource requests and limits explained: scheduling, throttling, and OOMKill

1
Comments 2
13 min read
Planning network checks before running them: a local-first workflow pattern

Planning network checks before running them: a local-first workflow pattern

1
Comments 6
4 min read
How to Optimize MongoDB on Bare Metal Servers: SRE Playbook

How to Optimize MongoDB on Bare Metal Servers: SRE Playbook

Comments
5 min read
Rename a Kubernetes PVC Without Losing Your Data: PersistentVolume Rebinding

Rename a Kubernetes PVC Without Losing Your Data: PersistentVolume Rebinding

Comments
4 min read
Performance Tuning: The Day the Server Got “Tired” and Started Acting Funny

Performance Tuning: The Day the Server Got “Tired” and Started Acting Funny

Comments
3 min read
Hiring SREs: What I Look For After Interviewing 100+ Candidates

Hiring SREs: What I Look For After Interviewing 100+ Candidates

Comments
3 min read
AI as an SRE Intern: What I Let Agents Touch During Incident Response

AI as an SRE Intern: What I Let Agents Touch During Incident Response

Comments
2 min read
Remetric: find waste in self-hosted Prometheus, Grafana, and Loki

Remetric: find waste in self-hosted Prometheus, Grafana, and Loki

Comments
6 min read
The Degradation Ladder: How Systems Fail Before They Fail

The Degradation Ladder: How Systems Fail Before They Fail

Comments
5 min read
Log Management at Scale: How We Cut Costs 70% Without Losing Signal

Log Management at Scale: How We Cut Costs 70% Without Losing Signal

Comments
2 min read
LiveOps Rollback Planning: What to Do When a Game Event Goes Wrong

LiveOps Rollback Planning: What to Do When a Game Event Goes Wrong

2
Comments 2
7 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.