Site Reliability Engineering Overview
SRE Principles
- Embracing risk: 100% reliability is neither possible nor desirable
- Service Level Objectives: Define what “good enough” means
- Eliminating toil: Automate repetitive manual work
- Monitoring distributed systems: Observability over monitoring
- Release engineering: Safe, automated deployments
- Simplicity: Simpler systems are more reliable
SRE vs DevOps
| Aspect | SRE | DevOps |
|---|---|---|
| Origin | Community-driven | |
| Focus | Reliability + automation | Culture + collaboration |
| Metrics | SLIs, SLOs, error budgets | DORA metrics |
| Practices | Toil reduction, chaos engineering | CI/CD, IaC |
| Overlap | Both: automation, monitoring, blameless culture |