What does the observability tool landscape look like, and what are the pieces you end up visualising?
Open-source building blocks (Prometheus, Grafana, Jaeger, OpenTelemetry) versus integrated commercial platforms (Dynatrace, Datadog, New Relic, Splunk, Azure Monitor) — and the outputs are dashboards, traces, service maps and alerting.
Open source is assembled from parts: Prometheus for metrics, Jaeger or Tempo for traces, Loki or Elasticsearch for logs, Grafana for visualisation, OpenTelemetry to produce and ship the data. You own the integration and the operation; you also own the bill and the data.
Commercial platforms deliver the whole chain as a product, with the correlation between the pillars already wired up and increasingly with machine-learning-assisted analysis. You pay per host or per ingested gigabyte — and since observability data volume grows with the estate, that bill is the main thing to model in advance.
Either way the outputs are the same four:
- Dashboards — the state of a system at a glance.
- Traces — an individual request's path, usually drawn as a waterfall of spans.
- Service maps — the dependency graph between services, derived from the traces rather than from documentation, which is what keeps it accurate.
- Alerting — the push channel to a human.
Because instrumentation via OpenTelemetry is portable, the open-source-versus-commercial decision is no longer irreversible — which is exactly the point of standardising on it first.
Go deeper:
Modern Observability with the LGTM stack (GPN21) — 30 minutes on assembling Loki, Grafana, Tempo and Mimir around OpenTelemetry.
Grafana Tempo documentation — the open-source trace back end in the table.
CNCF Landscape guide — the map of the whole cloud-native ecosystem, observability included.