⚡THE SHORT ANSWER
Alert fatigue is an existential threat to reliability: when an on-call rotation receives dozens of nocturnal pages for non-critical warnings (e.g. 'CPU at 82%', 'Cron job warning', 'Single disk at 75%'), engineers become desensitized and inevitably sleep through true P0 outages. Google SRE and Rob Ewaschuk's alerting doctrine defines the golden rule: Every alert that pages a human MUST represent an urgent, actionable problem affecting real users or burning error budgets, and MUST link directly to a verified runbook. If an alert requires no human action within 15 minutes, it must NOT page—it should be routed to a daily Slack digest or Jira queue. Enforcing an 'alert deletion policy' for noisy alarms restores on-call sanity and slashes incident response latency.
Engineering Handbook & Failure Dynamics
6-Dimensional Architecture Breakdown⚙️1. Underlying Mechanism
Execution🎯2. Appropriate Use Context
Scope⚠️3. Production Failure Modes
P0 Risk📡4. Diagnostic Signals & Telemetry
Telemetry🛡️5. Prevention & Safeguards
Safeguards⚖️6. Architectural Trade-offs
Trade-offCase Study (TinyCTO In-Field Example)
An e-commerce platform's on-call engineers received 120 pages/week across 80 microservices. The team implemented symptom-based alerting: deleting all 400 CPU/memory threshold alerts and replacing them with 6 multi-window error-budget burn rate alerts on user-facing API routes with mandatory runbook links. Weekly pages dropped from 120 to 4, on-call engineer retention stabilized, and MTTR dropped by 55% because engineers immediately trusted and acted on every page.
Interactive Concept Drills
2 CardsWhat is the 'Golden Rule' of on-call paging alerts?
Why should teams alert on symptoms (SLOs) rather than causes (CPU/Memory)?
On-Call Alert Fatigue & Actionable Paging Threshold Hygiene — Technical FAQ
What is a healthy target for Pager Load per on-call shift?
Google SRE guidelines recommend a maximum of 2 actionable incidents per 24-hour shift to ensure engineers have time to recover and complete postmortems.
What should happen to an alert that fired 10 times in a week with zero human action required?
It must be deleted immediately or converted to an asynchronous daily report. It has zero business waking engineers.
🤖 AEO & Key Facts Summary
Key Architectural Facts
- ▸
Every page must be urgent, actionable, linked to a runbook, and impact user SLOs.
- ▸
Alert on symptoms (user errors/latency) rather than causes (raw CPU/memory).
- ▸
Target pager load is <2 pages per 24-hour shift.
- ▸
Noisy, non-actionable alarms must be aggressively deleted during weekly handoffs.
Common Misconceptions
- ✗
Misconception: More alerts mean a safer system (False: Too many alerts guarantee alert fatigue and missed outages).
- ✗
Misconception: 90% CPU utilization always requires waking up an engineer (False: If user requests succeed within SLA, high CPU is healthy efficiency).
Decision & Governance Guidance
Audit all PagerDuty alerts to enforce mandatory runbook links. Delete any alert that has triggered without requiring manual action.
Authoritative Sources & Standards
- [OFFICIAL_DOCUMENTATION]My Philosophy on Alerting: SRE Principles and Alert Fatigue— Rob Ewaschuk / Google SRE
