Which categories of tooling make up a monitoring stack, and why is a time-series database used instead of a normal database?
Collection agents, a monitoring server, a time-series database, a rules/analysis engine, visualisation and alerting — and metrics are stored in a TSDB because they are append-only, timestamped and queried by range, which a TSDB compresses and scans far better than a relational store.
| Job | Tools |
|---|---|
| Collection (agents / exporters) | Prometheus Node Exporter, Windows Exporter, Zabbix Agent, Datadog Agent, Telegraf, SNMP agents, OpenTelemetry Collector |
| Monitoring server | Prometheus, Zabbix Server, PRTG Core Server, Nagios, Icinga, Datadog, Azure Monitor |
| Storage (time-series database) | Prometheus TSDB, InfluxDB, VictoriaMetrics, TimescaleDB, OpenTSDB |
| Analysis and rules | Prometheus Alertmanager, Zabbix Triggers, Nagios checks, Datadog Monitors, Azure Monitor Alerts |
| Visualisation | Grafana, Zabbix Dashboard, PRTG Dashboard, Datadog Dashboard, Azure Dashboards |
| Alerting / notification | Alertmanager, Grafana Alerting, Zabbix Actions, PagerDuty, OpsGenie, ServiceNow |
Why a TSDB. A monitoring system writes the same handful of metric names millions of times a day, never updates a past value, and asks questions of the form "CPU for these hosts over the last six hours, averaged per minute". That workload is nothing like a business database: values arrive in timestamp order and change slowly, so a TSDB stores them with delta and compression schemes that shrink them enormously, and indexes by time and labels so a range scan is cheap. Downsampling old data (per-second detail for a day, per-hour averages for a year) is built in, which is how a year of history stays affordable.
Go deeper:
Prometheus — Storage — how a real TSDB lays out blocks, compresses samples and retains data.
Time series database — Wikipedia — why the workload is different and the main products.
Prometheus Alertmanager — grouping, inhibition and silencing: the rules/alerting layer in practice.