You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

ElasticSearch查询实现单词连续字符匹配及问题求助

Solution for Elasticsearch Continuous Substring Matching

Let's break down what's wrong with your current query first, then walk through the correct approach to meet your requirements.

What's Wrong with Your Existing Query?

Your current multi_match query has two key issues that lead to unwanted results and fail to match middle substrings like "ohn":

  1. phrase_prefix only matches prefixes: This type will only match the start of words, so it can't catch substrings like "ohn" from "John".
  2. operator: OR is too permissive: When you search "john x", the "john" part matches "John" in your index, and the OR operator means only one match is needed to return the document—hence the unwanted result.

Correct Approach: Use Ngram Tokenization

To match any continuous substring of words (prefix, middle, suffix) while ensuring all search terms have a match, we'll use an ngram analyzer at index time. This splits each word into all possible continuous character sequences, making substring searches efficient and accurate.

Step 1: Create Index with Ngram Analyzer

First, define an index with a custom ngram analyzer that splits words into substrings, and apply it to your fields:

PUT /your_index_name
{
  "settings": {
    "analysis": {
      "analyzer": {
        "substring_analyzer": {
          "tokenizer": "substring_tokenizer",
          "filter": ["lowercase"]
        }
      },
      "tokenizer": {
        "substring_tokenizer": {
          "type": "ngram",
          "min_gram": 1,  // Adjust to 2 if you don't need single-character matches
          "max_gram": 20  // Match the longest word length in your data
        }
      }
    }
  },
  "mappings": {
    "properties": {
      "Field1": {
        "type": "text",
        "analyzer": "substring_analyzer",
        "search_analyzer": "standard"  // Keep search terms as whole tokens
      },
      "Field2": {
        "type": "text",
        "analyzer": "substring_analyzer",
        "search_analyzer": "standard"
      }
    }
  }
}

Step 2: Search with AND Operator

Use a multi_match query with operator: AND to ensure every search term has a matching substring in your fields. This way, "john x" won't return "John Doe" because "x" has no matching substring:

{
  "query": {
    "multi_match": {
      "query": "ohn do",  // Replace with your search term (e.g., "n doe", "john")
      "operator": "AND",
      "fields": ["Field1", "Field2"]
    }
  }
}

Testing Your Cases

All your required search terms will return "John Doe":

  • john doe: Exact match of both words
  • john do: "john" matches "John", "do" matches the start of "Doe"
  • ohn do: "ohn" matches the middle substring of "John", "do" matches "Doe"
  • john: Matches "John"
  • n doe: "n" matches the end of "John", "doe" matches "Doe"

And john x will not return the result, since "x" has no corresponding substring in your document.

Why Not Use Wildcards?

Wildcards like *ohn* work but are inefficient for large datasets—Elasticsearch can't use the inverted index when a wildcard starts with *, forcing a full scan of documents. Ngram tokenization precomputes all substrings at index time, making searches fast and scalable.

内容的提问来源于stack exchange,提问作者user2748107

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:36:18