Mission Logs
They are monolithic, so they scale out badly, and they are event-based: they report the state of a host or service at one moment and fire when something breaks.
Both traits made sense when these tools were born. A data centre had a known, bounded size, and one central server coul...
Q What is PromQL, and why didn't Prometheus just use SQL?
PromQL is Prometheus's own query language. Its syntax resembles SQL, but it is built for time series, because several of the query types Prometheus needs can't be expressed sensibly in SQL.
SQL is designed around tables and rows. Prometheus data is series of timestamped values id...
Q Why does switching from the Node Exporter to Telegraf in a running setup break alerts and dashboards...
The two collectors name their metrics (and labels) differently, and every alert rule and every Grafana panel is a PromQL query that selects metrics by those names.
An alert like node_filesystem_avail_bytes < … or a panel graphing node_cpu_seconds_total simply finds no data once t...
Q The Prometheus systemd unit contains Wants=network-online.target, After=network-online.target and Wa...
Once the network is up and the system reaches the normal multi-user state, the regular boot target of a server, provided the unit has been enabled.
* Network first, then Prometheus, as part of the normal multi-user boot. *
Wants=network-online.target pulls in the "network is on...
Q Write PromQL queries for a Windows Server dashboard: CPU usage in percent, RAM usage, network traffi...
CPU: 100 minus the idle rate. RAM: total minus free. Network: rate() of the byte counters. Disk: size minus free per volume, or the fraction used.
* From a per-core idle counter to one busy-percent number per host. *
# CPU usage in %, averaged over all cores
100 - (avg by (insta...
Q What is "trending" in operations, and why can't a platform run without it?
Trending is collecting load metrics over time to forecast capacity, so you know when to order hardware before you run out.
New hardware is not available overnight. Choosing, ordering, delivering, cabling and commissioning servers takes weeks, in the worst case months. To order in...
Q What is an exporter, and what is "scraping"?
An exporter is a program on the target system that reads values (CPU, RAM, a database's internals) and serves them as metrics over an HTTP API. Scraping is Prometheus fetching that endpoint.
This is the biggest mental shift from classic monitoring. Nagios-style checks are scripts...
Q Every extra exporter opens another port on the host. What problem does that cause, and how can Teleg...
Each port must be allowed in the host firewall, which in a corporate setting costs compliance work and turns the firewall into Swiss cheese. Telegraf can scrape local exporters on 127.0.0.1 and re-expose everything through its single port.
* Three firewall holes versus one: Tele...
Q What does --web.listen-address=0.0.0.0:9090 do, and is it a sensible setting?
Prometheus listens on port 9090 on all network interfaces, so anyone who can reach the host can open the web UI and API. That's fine in an isolated lab but risky in production, because there is no authentication by default.
* 0.0.0.0 opens every interface; 127.0.0.1 keeps it loc...
Q Why does storing monitoring metrics in a relational database like MySQL become a bottleneck?
The write volume is huge, and reading a metric over a time range means pulling it out of a table row by row, which gets slow as history grows.
Writing: 4000 systems with 200 values each, polled every 15 seconds, is 800,000 values per poll round, more than 3 million rows a minute,...