如何构建高效正则表达式优化Elasticsearch含目标字符的查询性能
Great question—using .*s.* in regexp queries is definitely a performance trap in Elasticsearch, especially as your dataset grows. Let's walk through why that's slow and the better ways to get the same "contains letter s" result without the overhead:
Why .*s.* Is So Slow
Here's the thing: Elasticsearch's inverted index is built to quickly look up terms that start with specific characters. When you use a leading wildcard like .*, you're telling Elasticsearch it can't use that index advantage—it has to scan every single term in your name field to check if it matches the regex. That's a full scan, and it gets exponentially worse the more data you have.
Better Alternatives, Ranked by Performance
1. Use an Ngram Tokenizer (Best Long-Term Solution)
If you can reindex your data (and you should, if this is a common query), ngram tokenization is the way to go. It shifts the work to index time, so queries become lightning fast.
Step 1: Create the Index with a Custom Ngram Analyzer
First, drop your existing strings index (if you can), then set up a new one with an analyzer that breaks strings into small substrings (including single characters):
PUT strings { "settings": { "analysis": { "analyzer": { "ngram_analyzer": { "tokenizer": "ngram_tokenizer", "filter": ["lowercase"] } }, "tokenizer": { "ngram_tokenizer": { "type": "ngram", "min_gram": 1, # Index single characters like "s" "max_gram": 20 # Adjust based on your longest string length } } } }, "mappings": { "properties": { "name": { "type": "text", "analyzer": "ngram_analyzer", "fields": { "keyword": { # Keep keyword field for exact matches if needed "type": "keyword" } } } } } }
Step 2: Insert Your Docs (Same as Before)
POST strings/_doc/1 { "name": "hello world" } POST strings/_doc/2 { "name": "sunshine" } POST strings/_doc/3 { "name": "test string" }
Step 3: Query with a Simple match
Now you can find all strings containing "s" with a query that uses the inverted index efficiently:
GET strings/_search { "query": { "match": { "name": "s" } } }
This will return docs 2 and 3, and it's way faster than regex because Elasticsearch just looks up the token "s" in its index instead of scanning every term.
2. Use a wildcard Query (Quick Fix, Better Than Regex)
If reindexing isn't an option right now, a wildcard query is a slight improvement over your current regex. Elasticsearch optimizes wildcard queries a bit more than arbitrary regex patterns:
GET strings/_search { "query": { "wildcard": { "name.keyword": "*s*" } } }
Note: Use the name.keyword sub-field here if your name is a text field—this ensures you're matching against the raw string, not analyzed tokens. This is still not great for large datasets (thanks to the leading *), but it's better than .*s.*.
3. Script-Based Query (Last Resort)
If you can't reindex and wildcard is still too slow, you can use a script to check for the presence of "s". This is still a scan, but it's more efficient than regex for simple contains checks:
GET strings/_search { "query": { "bool": { "filter": { "script": { "script": "doc['name.keyword'].value.contains('s')" } } } } }
Again, use the keyword sub-field here to avoid dealing with analyzed text. This works but isn't scalable for huge datasets—reindexing with ngrams is still better.
Quick Recap
- Avoid leading wildcards/regex like
.*s.*—they force full index scans. - Ngram tokenization is the best bet for repeated "contains" queries—do the work once at index time, not every time you query.
- Wildcard/script queries are temporary fixes if reindexing isn't possible.
内容的提问来源于stack exchange,提问作者K.Wu

