What are Ranum's three laws of log analysis, and how can the first and the third be reconciled?
Never keep more than you could conceivably look at; the number of times an uninteresting thing happens is itself interesting; keep everything you possibly can — except where that conflicts with the first law.
* Laws 1 and 3 pull in opposite directions; tiering keeps everything without indexing everything. *
Marcus Ranum's laws, formulated in 2004 and still the sanest guidance on the subject:
- "Never keep more than you can conceive of possibly looking at." Data you will never examine is not an asset — it is storage cost, search latency and, if it contains personal data, liability.
- "The number of times an uninteresting thing happens is an interesting thing." A single failed login is noise; four thousand in a minute is an attack. The individual entry is worthless, the rate is the signal — which is why aggregation, not reading, is the primary mode of log analysis.
- "Keep everything you possibly can, except where you come into conflict with the First Law." You cannot know in advance which field the next investigation will need, so discard reluctantly.
Laws 1 and 3 look contradictory and the tension is deliberate. The resolution in practice is tiering: keep everything, but not everywhere — a short, fully indexed hot window you actually search, and a cheap cold archive for the rest, aggregating the high-volume, low-value classes into counts rather than storing each line.
Which leads to the philosophy behind the whole exercise: logs are just data; processed and analysed, they become information. Collection alone produces nothing. Parsing, normalising, correlating and aggregating are where the value is created.
Go deeper:
Marcus Ranum — System Logging and Log Analysis (tutorial slides, PDF) — the laws in their author's own hands-on course, with the reasoning behind each.