LOGBOOK

HELP

Quiz Entry - updated: 2026.10.01

Prometheus is poor at keeping years of metrics. How is long-term trending usually solved, and what is downsampling?

Prometheus keeps recent data and writes selected metrics to a second TSDB built for long-term storage, such as InfluxDB or Graphite. Downsampling means averaging data into coarser intervals before storing it long term.

A day of vCPU usage sampled every 15 seconds as a noisy band, with hourly averages drawn as a step line that keeps the same trend

* Same trend, 240 times less data: hourly averages instead of 15-second samples. *

Prometheus gets slow and sluggish when fed historical data for hundreds of hosts going back years. The fix is a division of labour:

  1. Prometheus uses a storage plug-in (remote write) to send selected metrics to InfluxDB, Graphite, OpenTSDB or Gnocchi.
  2. Downsample first. Nobody needs a load balancer's 15-second response times from a Saturday three years ago. What matters long term is, say, vCPU and RAM usage per hour or day. Averaging to larger intervals cuts the data volume drastically.
  3. Data that reached the long-term store is deleted from Prometheus.
  4. Grafana gets the long-term store as an additional data source, with its own dashboards.

Alerting stays with Prometheus only, so the long-term store is not connected to the Alertmanager.

Go deeper:

From Quiz: ITIA / Monitoring Lab: Prometheus and Grafana | Updated: Oct 01, 2026