What is the difference between analyzed and not-analyzed fields in Elasticsearch, and why does it change how you search?
An analyzed (text) field is split into lower-cased individual terms, so you can search for any word in it; a not-analyzed (keyword, often .raw) field is stored as one exact string, so it matches only as a whole, case-sensitively.
* Search words in the analyzed field, exact values in the keyword field. *
Take the string Set the shape to semi-transparent by calling set_trans(5):
- Analyzed: stored as the terms
set, the, shape, to, semi, transparent, by, calling, set_trans, 5. A search fortransparentfinds it. - Not analyzed: stored as the single exact string. Only the complete phrase, with matching case, finds it.
Analyzed (text) |
Not analyzed (keyword / .raw) |
|
|---|---|---|
| Good for | Free-text search in messages | Exact values: hostnames, paths, IDs, status codes |
| Search | Any part of it | Whole value, case-sensitive |
| Aggregations and sorting | Not practical | Yes |
That is why many fields exist twice, e.g. parser and parser.raw (or message and message.keyword): the first for searching, the second for exact matches, counting and top-N charts. If an exact search mysteriously returns nothing, you are probably querying the analyzed version, or the case does not match.
Go deeper:
Elastic — Text analysis — tokenizers, filters and analyzers in detail.
Elastic — Keyword type family — the not-analyzed field type for exact values and aggregations.