Quiz Entry - updated: 2026.10.01
What two traits of classic monitoring tools like Nagios and Zabbix make them a poor fit for large cloud platforms?
They are monolithic, so they scale out badly, and they are event-based: they report the state of a host or service at one moment and fire when something breaks.
Both traits made sense when these tools were born. A data centre had a known, bounded size, and one central server could poll every host. "Is this service up right now?" was the question that mattered, so the design centres on checks that return a status and trigger a chain of events ending in the on-call admin's phone ringing.
A cloud or Kubernetes platform breaks both assumptions:
- Size is open-ended. Thousands of hosts that come and go, so one monolithic server becomes the bottleneck.
- A single failure is often not an emergency. What matters is whether the interface to the user still works. If one of ten API instances behind a load balancer dies, nobody needs to be woken at 3 a.m.; it can be fixed the next day.
- Point-in-time state is not enough. The operator also has to know when to scale, which needs history, not just the current status.
Go deeper:
Nagios — Wikipedia — the classic check-based monitor Prometheus is usually contrasted with.
Google SRE book — Monitoring Distributed Systems — why large platforms alert on user-facing symptoms, not every broken instance.