Jaeger's trace search lets you filter by service, operation, attributes, time window and minimum/maximum duration. How do you use those filters to find the one trace that explains a problem?
Narrow by the symptom: restrict to the affected service and time window, then filter on the attribute or duration that describes the failure, such as error=true or a minimum duration of 1 s.
Even a small demo produces hundreds of traces per minute, so browsing is hopeless. Each filter encodes part of the complaint:
| Complaint | Filter |
|---|---|
| "Checkout fails since about 10 minutes ago" | Service checkout, lookback 15 min, attributes error=true |
| "Some product pages are very slow" | Service frontend, operation GET /api/products/..., min duration 1s |
| "Only requests that got an HTTP 500" | Attributes http.status_code=500 |
The result list shows each trace with its span count, the services it touched and its duration, and the scatter plot above it plots duration against time. The outliers at the top of that plot are the interesting ones. Selecting two traces and pressing Compare shows the structural difference between a slow request and a normal one, which often points directly at the extra or slower span.