What makes unplanned downtime so expensive, and why has finding the cause of an outage become harder rather than easier?
Direct cost is only part of it — the reputational damage often outweighs the lost revenue — and in today's hybrid, SaaS-dependent estates fewer than 40% of IT leaders can pinpoint an outage's cause immediately.
Industry surveys consistently put the cost of large-enterprise downtime in the hundreds of millions per year across a sector, but the number that should worry an engineer is a different one: only about 38% of IT leaders can identify the root cause of an outage straight away and correctly. The rest are guessing, or waiting.
The reason is structural. A failing transaction today crosses your own systems, someone else's cloud, and two or three third-party SaaS platforms you cannot log into. Each piece may look healthy in isolation while the transaction as a whole fails, and you have no visibility into most of the path.
That gap is exactly what organisations are buying when they invest in observability platforms (increasingly with machine-learning-assisted analysis): not more dashboards, but the ability to answer which component caused the loss, in minutes rather than days. Reputational damage scales with how long the answer takes.
Go deeper:
heise: IT-Ausfälle immer teurer für grosse Unternehmen — the news write-up of the downtime-cost study (German).
Splunk: Der 600-Milliarden-Dollar-Weckruf (press release) — the study itself, including the root-cause-identification figure (German).