What makes an alerting setup useful rather than something people learn to ignore?
Thresholds that fit the metric (static or anomaly-based), an escalation chain that reaches a human who can act, and deliberate work against alert fatigue — bundling, prioritising and de-duplicating.
* An escalation chain: each rung fires only when the previous one goes unanswered. *
Thresholds come in two flavours. Static ones are simple and predictable — CPU above 90% for five minutes — but they are wrong for anything with a daily or seasonal rhythm. Dynamic or anomaly-based ones learn the normal shape of the metric and fire on deviation, which catches "traffic is a third of what Tuesday morning usually looks like" — a serious incident that no static threshold would ever notice.
Escalation chains exist because an alert nobody acknowledges is not an alert: email → chat → SMS → on-call phone, each step triggered when the previous one goes unanswered.
Alert fatigue is the failure mode that quietly disables everything above. When a team receives hundreds of alerts a day, it stops reading them, and the one that mattered is lost in the pile. The three countermeasures:
- Bundle — one notification for a hundred hosts behind the same failed switch, not a hundred notifications.
- Prioritise — severity that reflects business impact, so the page at 3 a.m. is genuinely worth waking up for.
- De-duplicate — one open incident per underlying problem, no matter how many checks observe it.
Tip: the health metric for an alerting system is not how much it catches but what fraction of alerts led to an action. If it is low, the system is training its users to ignore it.
Go deeper:
Google SRE Book — Practical Alerting — how alert rules are built on time-series data and why symptom-based alerts beat cause-based ones.
Alarm fatigue — Wikipedia — the phenomenon as first studied in hospitals, with the same countermeasures.