LOGBOOK

HELP

Quiz Entry - updated: 2026.10.01

Read this alert rule. When does it fire, and what exactly is an "instance" here?

- alert: InstanceDown
  expr: up == 0
  for: 5m
  labels:
    severity: page

It fires when a scrape target has been unreachable for more than five minutes. An "instance" is one exporter endpoint, not a whole machine.

Timeline of up for one target: the scrape fails at minute 3, the alert is pending for 5 minutes, fires at minute 8 and resolves when the target returns at minute 11

* for: 5m in action: pending first, firing only after five minutes of up == 0. *

  • up is a metric Prometheus creates itself for every target: 1 if the last scrape succeeded, 0 if it failed.
  • expr: up == 0 selects every target whose last scrape failed.
  • for: 5m means the condition must hold for five minutes before the alert fires. Before that the alert is only pending, which avoids pages for one dropped scrape.
  • labels: severity: page attaches a label the Alertmanager can route on (e.g. page the on-call person).
  • Annotations such as summary: "Instance {{ $labels.instance }} down" use templating, so one rule produces a readable message for whichever target failed.

Because an instance is an exporter, a host that dies completely raises one alert per exporter running on it.

Go deeper:

From Quiz: ITIA / Monitoring Lab: Prometheus and Grafana | Updated: Oct 01, 2026