You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按词匹配所有子串?分词器需求与查询优化问询

Solution for Matching Exact Contiguous Substrings & Generating All Substrings

1. Fixing the Extra Matches in Your Query

The problem with your current match query is that it’s pulling in terms that share individual words with your input (like "sell car" shares "sell" and "and me" shares "and") but aren’t actual contiguous substrings of the input text. Here are two straightforward ways to fix this:

Approach 1: Query + Client-Side Filtering

This is the simplest fix if you want to keep using your existing query structure:

  1. Run your match query to get all candidate terms that overlap with your input’s words.
  2. Filter the results on the client side to only keep terms that are exact contiguous substrings of your input.

Example Code (Python)

from elasticsearch import Elasticsearch

# Initialize your Elasticsearch client
es = Elasticsearch(["http://your-es-host:9200"])

input_text = "sell everything and live in".lower()
initial_query = {
    "match": {
        "text": {
            "query": input_text
        }
    }
}

# Fetch initial candidate results
response = es.search(index="your-index-name", query=initial_query, size=100)
initial_results = [hit["_source"]["text"] for hit in response["hits"]["hits"]]

# Filter to keep only valid contiguous substrings
filtered_results = [term for term in initial_results if term.lower() in input_text]

print(filtered_results)
# Output: ["sell", "everything", "and", "live", "in", "live in"]

Approach 2: Pre-Generate Substrings + Terms Query

If you want to avoid client-side filtering entirely, generate all possible contiguous substrings from your input first, then use a terms query to fetch only those exact terms from your index. This ensures you never get extra matches in the first place:

def generate_contiguous_substrings(text):
    words = text.lower().split()
    substrings = []
    n = len(words)
    for i in range(n):
        for j in range(i+1, n+1):
            substrings.append(' '.join(words[i:j]))
    return substrings

# Generate all substrings from your input
input_substrings = generate_contiguous_substrings(input_text)

# Build a terms query for exact matches (uses keyword field for precision)
terms_query = {
    "terms": {
        "text.keyword": input_substrings
    }
}

# Fetch exact matches from the index
response = es.search(index="your-index-name", query=terms_query, size=100)
final_results = [hit["_source"]["text"] for hit in response["hits"]["hits"]]

print(final_results)
# Output: ["sell", "everything", "and", "live", "in", "live in"]

Note: This requires your text field to have a keyword sub-field (standard in most Elasticsearch mappings) to enable exact term matching.

2. Implementing the Substring Tokenizer

To generate every possible contiguous word substring from a text (like turning "So do I" into ["so", "do", "i", "so do", "do i", "so do i"]), use this simple function:

Python Implementation

def generate_all_contiguous_substrings(text):
    # Clean input and split into words
    words = text.strip().lower().split()
    substrings = []
    num_words = len(words)
    
    # Generate all contiguous word sequences
    for start in range(num_words):
        for end in range(start + 1, num_words + 1):
            substring = ' '.join(words[start:end])
            substrings.append(substring)
    
    return substrings

# Test with your example input
print(generate_all_contiguous_substrings("So do I"))
# Output: ["so", "so do", "so do i", "do", "do i", "i"]

If you want the output ordered like your example (single words first, then longer phrases), sort the results by word count then alphabetically:

sorted_substrings = sorted(substrings, key=lambda x: (len(x.split()), x))
print(sorted_substrings)
# Output: ["so", "do", "i", "so do", "do i", "so do i"]

This function works for any input text, generating every possible contiguous word combination needed for your use case.

内容的提问来源于stack exchange,提问作者m1neral

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:28:15