如何在Elasticsearch中对英文连字符词进行索引与搜索?
Great question—this is a super common pain point with hyphenated terms in Elasticsearch, and you absolutely can solve this with analyzer configurations instead of messy application-layer hacks or manual synonym lists. Here's a robust approach that handles both scenarios you mentioned:
word_delimiter_graph Filter The core idea is to build a custom analyzer that automatically handles two critical matching needs:
- Hyphenated terms ↔ their non-hyphenated equivalents (e.g.,
pre-eclampsia↔preeclampsia) - Space-searched phrases ↔ hyphenated terms (e.g.,
jean claude↔jean-claude)
We’ll use Elasticsearch’s word_delimiter_graph filter (available in 7.6+), which is purpose-built for splitting and combining terms with separators like hyphens while preserving phrase matching accuracy.
Step 1: Create the Index with Custom Analyzer
First, define your index with the custom analyzer, including the specialized filter for hyphen handling:
PUT /hyphenated_docs { "settings": { "analysis": { "analyzer": { "hyphen_smart_analyzer": { "tokenizer": "standard", "filter": [ "hyphen_word_delimiter", "lowercase" ] } }, "filter": { "hyphen_word_delimiter": { "type": "word_delimiter_graph", "preserve_original": true, // Keeps the original hyphenated term (e.g., "pre-eclampsia") "catenate_words": true, // Combines split word parts into a single term (e.g., "pre" + "eclampsia" → "preeclampsia") "split_on_hyphens": true, // Splits terms at hyphens (explicit for clarity) "generate_word_parts": true // Generates individual word parts (e.g., "pre", "eclampsia") } } } }, "mappings": { "properties": { "content": { "type": "text", "analyzer": "hyphen_smart_analyzer", "search_analyzer": "hyphen_smart_analyzer" // Use the same logic for indexing and searching } } } }
Step 2: How It Solves Your Scenarios
Let’s break down the behavior for your key use cases:
Case 1: Hyphenated ↔ Non-Hyphenated Terms
When indexing pre-eclampsia, the analyzer generates these tokens:
pre-eclampsia(original term, preserved)pre(split word part)eclampsia(split word part)preeclampsia(combined word parts)
When indexing preeclampsia, the standard tokenizer generates just preeclampsia.
Now, searching for either term will match both documents:
- Searching
pre-eclampsiaincludes thepreeclampsiatoken in its query, matching the non-hyphenated document. - Searching
preeclampsiamatches the combined token from the hyphenated document.
Case 2: Space-Searched Phrases ↔ Hyphenated Terms
When indexing jean-claude, the analyzer generates:
jean-claude(original)jean(split part)claude(split part)jeanclaude(combined)
When searching jean claude, the query splits into jean and claude tokens—both of which exist in the indexed tokens for jean-claude, so the document matches.
Conversely, searching jean-claude will also match documents with jean claude, since the query generates jean and claude tokens that align with the space-separated text.
Step 3: Verify the Analyzer
You can test the analyzer’s output directly with the _analyze endpoint:
POST /hyphenated_docs/_analyze { "analyzer": "hyphen_smart_analyzer", "text": "happy-go-lucky" }
This returns tokens: happy-go-lucky, happy, go, lucky, happygolucky—covering all possible matching scenarios for that term.
Key Advantages Over Application-Layer Workarounds
- No manual synonym lists: Works for all hyphenated terms automatically, even with large, diverse datasets.
- Preserves phrase accuracy: The
word_delimiter_graphfilter maintains term position information, so phrase queries (e.g.,"happy go lucky") still work correctly. - Single source of truth: All logic lives in Elasticsearch, avoiding inconsistencies between application code and search behavior.
内容的提问来源于stack exchange,提问作者michael

