Question
What two traits of classic monitoring tools like Nagios and Zabbix make them a poor fit for large cloud platforms?
Answer
They are monolithic, so they scale out badly, and they are event-based: they report the state of a host or service at one moment and fire when something breaks.
Both traits made sense when these tools were born. A data centre had a known, bounded size, and one central server could poll every host. "Is this service up right now?" was the question that mattered, so the design centres on checks that return a status and trigger a chain of events ending in the on-call admin's phone ringing.
A cloud or Kubernetes platform breaks both assumptions:
- Size is open-ended. Thousands of hosts that come and go, so one monolithic server becomes the bottleneck.
- A single failure is often not an emergency. What matters is whether the interface to the user still works. If one of ten API instances behind a load balancer dies, nobody needs to be woken at 3 a.m.; it can be fixed the next day.
- Point-in-time state is not enough. The operator also has to know when to scale, which needs history, not just the current status.
Go deeper:
Nagios — Wikipedia — the classic check-based monitor Prometheus is usually contrasted with.
Google SRE book — Monitoring Distributed Systems — why large platforms alert on user-facing symptoms, not every broken instance.
Note saved — thanks!
Question
What is "trending" in operations, and why can't a platform run without it?
Answer
Trending is collecting load metrics over time to forecast capacity, so you know when to order hardware before you run out.
New hardware is not available overnight. Choosing, ordering, delivering, cabling and commissioning servers takes weeks, in the worst case months. To order in time, an operator has to measure the platform's utilisation continuously and check how much headroom is left, then extrapolate from the growth of recent months.
So trending answers a different question from monitoring:
| Monitoring | Trending | |
|---|---|---|
| Question | Is it healthy now? | Where is it heading? |
| Data needed | Current value | Long series of values |
| Result | An alert | A forecast (capacity plan) |
Together with alerting, the three form MAT: Monitoring, Alerting and Trending.
Go deeper:
Capacity planning — Wikipedia — the discipline trending data feeds.
Note saved — thanks!