ITIA Logs
The write volume is huge, and reading a metric over a time range means pulling it out of a table row by row, which gets slow as history grows.
Writing: 4000 systems with 200 values each, polled every 15 seconds, is 800,000 values per poll round, more than 3 million rows a minute,...
Q When do you need the Push-Gateway, and how does it fit the pull model?
For sources Prometheus can't scrape directly, such as network gear reachable only via SNMP or short-lived batch jobs. They push their metrics to the gateway, and Prometheus scrapes the gateway like any other exporter.
* Pushers on one side, a normal scrape on the other. *
Pull o...
Q Prometheus has no built-in clustering. How do you scale it to more targets than one instance can han...
By sharding: several independent Prometheus instances each scrape a different part of the estate. For high availability, let two instances scrape the same part, so the data exists twice independently.
* Shards for scale, a second scraper per shard for HA, one place to administer...
Q What is the /metrics endpoint of an exporter, and how is a Prometheus metric built?
/metrics is the HTTP page Prometheus scrapes. It lists every metric as plain text: a name, optional labels in braces and the current value, with # HELP and # TYPE lines describing it.
* Name, labels, value, type: the four parts of every metric. *
An excerpt from a Node Exporter:...
Q What is a time series database (TSDB), and why is it faster than a relational database for metrics?
Its root structure is a timeline rather than tables of rows: every value is attached to a point in time, so "give me metric X from t1 to t2" is the native query and needs no transposing.
In a relational database, a time range has to be assembled: find the rows, sort them by times...
Q Who decides that something is an alert: Prometheus or the Alertmanager? What does the Alertmanager d...
Prometheus decides. Alert rules are defined and evaluated in Prometheus, which sends firing alerts to the Alertmanager. The Alertmanager only handles delivery: grouping, silencing, routing and paging.
* Prometheus decides, the Alertmanager delivers. *
The split of duties:
Pro...
Q Prometheus is poor at keeping years of metrics. How is long-term trending usually solved, and what i...
Prometheus keeps recent data and writes selected metrics to a second TSDB built for long-term storage, such as InfluxDB or Graphite. Downsampling means averaging data into coarser intervals before storing it long term.
* Same trend, 240 times less data: hourly averages instead o...
Q What is the difference between a counter and a gauge, and why does that matter for queries?
A counter only goes up (it resets to 0 on restart), like total bytes received. A gauge can go up and down, like current RAM in use. You almost always wrap a counter in rate(); a gauge you read directly.
* A counter only climbs (until a restart resets it); a gauge moves both ways...
Q How can one time series database cover monitoring, alerting and trending, instead of running two sep...
The current state of a service is just a time series with a single (latest) element. Add a rule engine that fires when values cross a threshold and monitoring falls out as a by-product of collecting trending data.
* One collection run, two uses: range queries for trending, lates...
Q Read this alert rule. When does it fire, and what exactly is an "instance" here? - alert: InstanceDo...
It fires when a scrape target has been unreachable for more than five minutes. An "instance" is one exporter endpoint, not a whole machine.
* for: 5m in action: pending first, firing only after five minutes of up == 0. *
up is a metric Prometheus creates itself for every target...