You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何构建高效正则表达式优化Elasticsearch含目标字符的查询性能

Optimizing "Contains 's'" Queries in Elasticsearch

Great question—using .*s.* in regexp queries is definitely a performance trap in Elasticsearch, especially as your dataset grows. Let's walk through why that's slow and the better ways to get the same "contains letter s" result without the overhead:

Why .*s.* Is So Slow

Here's the thing: Elasticsearch's inverted index is built to quickly look up terms that start with specific characters. When you use a leading wildcard like .*, you're telling Elasticsearch it can't use that index advantage—it has to scan every single term in your name field to check if it matches the regex. That's a full scan, and it gets exponentially worse the more data you have.

Better Alternatives, Ranked by Performance

1. Use an Ngram Tokenizer (Best Long-Term Solution)

If you can reindex your data (and you should, if this is a common query), ngram tokenization is the way to go. It shifts the work to index time, so queries become lightning fast.

Step 1: Create the Index with a Custom Ngram Analyzer

First, drop your existing strings index (if you can), then set up a new one with an analyzer that breaks strings into small substrings (including single characters):

PUT strings
{
  "settings": {
    "analysis": {
      "analyzer": {
        "ngram_analyzer": {
          "tokenizer": "ngram_tokenizer",
          "filter": ["lowercase"]
        }
      },
      "tokenizer": {
        "ngram_tokenizer": {
          "type": "ngram",
          "min_gram": 1,  # Index single characters like "s"
          "max_gram": 20  # Adjust based on your longest string length
        }
      }
    }
  },
  "mappings": {
    "properties": {
      "name": {
        "type": "text",
        "analyzer": "ngram_analyzer",
        "fields": {
          "keyword": {  # Keep keyword field for exact matches if needed
            "type": "keyword"
          }
        }
      }
    }
  }
}

Step 2: Insert Your Docs (Same as Before)

POST strings/_doc/1
{ "name": "hello world" }

POST strings/_doc/2
{ "name": "sunshine" }

POST strings/_doc/3
{ "name": "test string" }

Step 3: Query with a Simple match

Now you can find all strings containing "s" with a query that uses the inverted index efficiently:

GET strings/_search
{
  "query": {
    "match": {
      "name": "s"
    }
  }
}

This will return docs 2 and 3, and it's way faster than regex because Elasticsearch just looks up the token "s" in its index instead of scanning every term.

2. Use a wildcard Query (Quick Fix, Better Than Regex)

If reindexing isn't an option right now, a wildcard query is a slight improvement over your current regex. Elasticsearch optimizes wildcard queries a bit more than arbitrary regex patterns:

GET strings/_search
{
  "query": {
    "wildcard": {
      "name.keyword": "*s*"
    }
  }
}

Note: Use the name.keyword sub-field here if your name is a text field—this ensures you're matching against the raw string, not analyzed tokens. This is still not great for large datasets (thanks to the leading *), but it's better than .*s.*.

3. Script-Based Query (Last Resort)

If you can't reindex and wildcard is still too slow, you can use a script to check for the presence of "s". This is still a scan, but it's more efficient than regex for simple contains checks:

GET strings/_search
{
  "query": {
    "bool": {
      "filter": {
        "script": {
          "script": "doc['name.keyword'].value.contains('s')"
        }
      }
    }
  }
}

Again, use the keyword sub-field here to avoid dealing with analyzed text. This works but isn't scalable for huge datasets—reindexing with ngrams is still better.

Quick Recap

  • Avoid leading wildcards/regex like .*s.*—they force full index scans.
  • Ngram tokenization is the best bet for repeated "contains" queries—do the work once at index time, not every time you query.
  • Wildcard/script queries are temporary fixes if reindexing isn't possible.

内容的提问来源于stack exchange,提问作者K.Wu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:36:31