The user-visible response time of a distributed application is produced by a dozen systems together — how can monitoring attribute it to the individual components?
By tagging each request with an identifier at the point of entry and propagating it through every component, so each step reports its own share of the total and the end-to-end path can be reassembled afterwards.
* One request, one ID, five tiers: each reports its own timing, the collector reassembles the path — and 3.4 s of the 3.46 s belong to a single tier. *
The problem first. Business cares about one thing: did the transaction work, and how long did it take? But the transaction is served by a distributed system — browser, web server, application tiers in Java or .NET, a database, some external services. Monitoring each of those in isolation tells you every component is within its own thresholds, and still leaves you unable to say which of them consumed the three seconds the user waited.
The solution is to make the request, not the machine, the unit of observation:
- An identifier is attached when the request enters the system.
- It is passed along with every downstream call — HTTP header, message property, database context.
- Every component records its own timing against that identifier.
- Afterwards, all the records sharing the identifier are assembled into the complete path of that one request, with the time each tier contributed.
Commercial application performance monitoring tools built this early — Dynatrace's PurePath stitches one request across browser, web server, Java, .NET and database tiers — and the same idea, standardised and vendor-neutral, is what distributed tracing in OpenTelemetry does today. Once you have it, "the system is slow" becomes "Service B spends 3.4 seconds waiting on one query", which is an actionable statement.
Go deeper:
Dynatrace — PurePath — the commercial per-request tracing described here, in the vendor's own words.
Google — Dapper, a Large-Scale Distributed Systems Tracing Infrastructure — the 2010 paper that defined trace IDs, spans and sampling for distributed systems.