Which KPIs tell you whether IT operations is actually getting better, and what does each one measure?
MTTD and MTTR measure how fast you notice and fix, SLA compliance rate measures whether you kept your promises, alert-to-acknowledgment measures whether anyone is listening, and incident volume trend measures whether the underlying problems are being removed.
* MTTD and MTTR laid on one incident: the user's outage is the sum of both, and nothing is repaired before it is detected. *
| KPI | What it measures | Example target |
|---|---|---|
| MTTD (Mean Time to Detect) | Time until an incident is noticed | < 5 minutes for critical systems |
| MTTR (Mean Time to Resolve) | Time until an incident is resolved | < 15–20 minutes |
| SLA compliance rate | Share of incidents resolved within the SLA | > 98% per quarter |
| Alert-to-acknowledgment rate | Share of alerts acknowledged in time | > 95% |
| Incident volume trend | Number of incidents over time | Falling quarter over quarter |
They divide neatly into two groups. MTTD and MTTR are speed, and they are where better tooling helps most directly — observability shortens MTTR because the cause is found rather than guessed. Incident volume trend is health: it is the only one of the five that improves when problems are genuinely eliminated rather than handled faster, which makes it the hardest to move and the most honest.
The alert-to-acknowledgment rate is the early-warning indicator for alert fatigue. When it starts sliding, the alerting system is losing its audience — and every other number on this list will follow it down.
Tip: MTTD = detect, MTTR = resolve. You cannot fix what you have not noticed, so MTTD is the term that bounds the whole incident.
Go deeper:
Mean time to repair — Wikipedia — MTTR and its relatives (MTBF, MTTF) and how they are computed.