Search with autocomplete on Elasticsearch

Created: April 2015 Updated: Sept. 29, 2026

from work

Python, Elasticsearch

Suggestions after two letters, tolerant of typos and missing Polish characters.

People search a large catalogue by names they only half remember, with typos and without Polish characters, and suggestions should pop up while they're still typing. A query goes out as they type, so the answer has to come back in a fraction of a second, and order matters too, because out of dozens of matches only the first few are visible.

Why not LIKE

The first thing that comes to mind is LIKE in the database, and for a small table with exact names that's enough. Not here: with a leading wildcard it can't use an index, and “zol” will never find “Żółty”. Trigrams in the database, e.g. pg_trgm in PostgreSQL, cope with typos, but prefix suggestions and ranking would have to be put together by hand. So a separate search engine, even though it's a second system you have to keep in sync with the database.

How it works

A separate Elasticsearch index. At indexing time names are cut into word prefixes (edge n-grams), lowercased and stripped of Polish characters, and the query goes through the same processing, just without the cutting. Popularity, meaning how often an item has been chosen, is added to relevance, and typos are caught by a query with fuzziness.

Most of the work is done by the index settings, one analyzer for indexing and another for searching:

json
PUT /products
{
  "settings": {
    "analysis": {
      "filter": { "prefix": { "type": "edge_ngram", "min_gram": 2, "max_gram": 15 } },
      "analyzer": {
        "autocomplete": { "tokenizer": "standard", "filter": ["lowercase", "asciifolding", "prefix"] },
        "autocomplete_search": { "tokenizer": "standard", "filter": ["lowercase", "asciifolding"] }
      }
    }
  },
  "mappings": {
    "product": {
      "properties": {
        "name": { "type": "string", "analyzer": "autocomplete", "search_analyzer": "autocomplete_search" },
        "popularity": { "type": "integer" }
      }
    }
  }
}

The database stays the source of truth, background jobs update the index, and every now and then we rebuild it from scratch. In the browser a request goes out only after a short pause in typing, not on every keystroke, so Elasticsearch doesn't get asked about every letter.

Machine-translated from Polish (original).