Which metrics do the Astronomy Shop's services produce, and how do they differ from what a trace tells you?
Request rate, error rate, response time, CPU and memory: aggregated numbers over time, exported via the Collector to Prometheus. They show that something changed; a trace shows where in one request it happened.
| Metric | What it reveals |
|---|---|
| Request rate | Load and user activity, traffic spikes |
| Error rate | Share of failing requests |
| Response time | How long requests take, usually as percentiles |
| CPU usage | Whether a service is compute-bound |
| Memory usage | Leaks, pressure before an out-of-memory kill |
The difference to traces is one of shape. A metric is a cheap number sampled every few seconds and stored as a time series, so it can be kept for months, charted, and alerted on ("error rate above 5% for 5 minutes"). But it is an aggregate: a p95 latency of 1.4 seconds says nothing about which request was slow or why. A trace is the opposite, rich detail about individual requests, too expensive to keep for everything forever.
In practice you use them in that order: a metric fires the alarm, then traces from the affected time window explain it.
Go deeper:
OpenTelemetry — Metrics — counters, gauges and histograms, the instruments behind these numbers.
Prometheus (software) — Wikipedia — the time-series database that stores the shop's metrics.