Elasticsearch:Shingle索引下如何优先匹配非相邻词组合的搜索结果
Great question! The core issue here is that standard 2-shingle (bigram) indexing only captures adjacent term pairs, so "submarine sinks ships" gets split into submarine sinks and sinks ships—missing the non-adjacent submarine ships pair you want to prioritize. Here are several practical, actionable ways to fix this, depending on whether you can adjust your index setup or prefer query-time tweaks:
1. Add a Skip-Gram Token Filter to Your Index Analysis
Most modern search engines (like Elasticsearch or Solr) support skip-gram filters that generate n-grams with gaps between terms. For your case, you can configure a filter to create bigrams with a skip of 1 (capturing both adjacent pairs and pairs with one term in between).
Example Configuration (Elasticsearch):
Define a custom analyzer in your index settings to include this filter:
{ "settings": { "analysis": { "analyzer": { "skip_bigram_analyzer": { "tokenizer": "standard", "filter": ["lowercase", "skip_bigram_filter"] } }, "filter": { "skip_bigram_filter": { "type": "skip_gram", "min_gram": 2, "max_gram": 2, "skip_count": 1 } } } }, "mappings": { "properties": { "content": { "type": "text", "analyzer": "skip_bigram_analyzer", "search_analyzer": "standard" // Keep search flexibility with standard tokenizer } } } }
This analyzer will index "submarine sinks ships" with tokens: submarine, sinks, ships, submarine sinks, sinks ships, and submarine ships. Now when you search for your target phrase, documents containing the submarine ships bigram will naturally get a higher relevance score.
2. Boost the "submarine ships" Phrase at Query Time
If reindexing isn't an option, adjust your search query to explicitly boost documents that match the non-adjacent phrase. Combine your original query with a boosted phrase query that allows for a single term gap.
Example Query (Elasticsearch):
{ "query": { "bool": { "must": [ { "match": { "content": "submarine sinks ships" } } ], "should": [ { "phrase": { "content": { "query": "submarine ships", "slop": 1, // Allow one term between the two target words "boost": 2.0 // Double the weight of this phrase match } } } ] } } }
The slop:1 parameter tells the engine to accept "submarine" followed by "ships" with one term in between, and the boost ensures these matches rank higher than those with only adjacent bigrams.
3. Create a Dedicated Field for Non-Adjacent Bigrams
For more granular control, index a separate field that only captures skip bigrams (like term1-term3, term2-term4). This lets you explicitly target these pairs without mixing them with adjacent shingles.
Example Setup:
- Keep your existing shingle analyzer for the main
contentfield. - Add a new field
content_skip_bigramsusing the skip-gram analyzer from Solution 1. - Boost matches in this dedicated field during search:
{ "query": { "bool": { "must": [ { "match": { "content": "submarine sinks ships" } } ], "should": [ { "match": { "content_skip_bigrams": "submarine ships" } } ] } } }
4. Manual Shingle Preprocessing (Brute-Force for Simple Cases)
If your search engine doesn't support skip-grams out of the box, preprocess your text to manually inject the desired non-adjacent bigrams before indexing. For example, append "submarine ships" to the original text (or store it in a separate field) when indexing "submarine sinks ships". This is a straightforward workaround for small-scale use cases.
内容的提问来源于stack exchange,提问作者Raja Rajan

