You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Elasticsearch中对英文连字符词进行索引与搜索?

Great question—this is a super common pain point with hyphenated terms in Elasticsearch, and you absolutely can solve this with analyzer configurations instead of messy application-layer hacks or manual synonym lists. Here's a robust approach that handles both scenarios you mentioned:

Solution: Custom Analyzer with word_delimiter_graph Filter

The core idea is to build a custom analyzer that automatically handles two critical matching needs:

  1. Hyphenated terms ↔ their non-hyphenated equivalents (e.g., pre-eclampsia ↔ preeclampsia)
  2. Space-searched phrases ↔ hyphenated terms (e.g., jean claude ↔ jean-claude)

We’ll use Elasticsearch’s word_delimiter_graph filter (available in 7.6+), which is purpose-built for splitting and combining terms with separators like hyphens while preserving phrase matching accuracy.

Step 1: Create the Index with Custom Analyzer

First, define your index with the custom analyzer, including the specialized filter for hyphen handling:

PUT /hyphenated_docs
{
  "settings": {
    "analysis": {
      "analyzer": {
        "hyphen_smart_analyzer": {
          "tokenizer": "standard",
          "filter": [
            "hyphen_word_delimiter",
            "lowercase"
          ]
        }
      },
      "filter": {
        "hyphen_word_delimiter": {
          "type": "word_delimiter_graph",
          "preserve_original": true,  // Keeps the original hyphenated term (e.g., "pre-eclampsia")
          "catenate_words": true,     // Combines split word parts into a single term (e.g., "pre" + "eclampsia" → "preeclampsia")
          "split_on_hyphens": true,   // Splits terms at hyphens (explicit for clarity)
          "generate_word_parts": true // Generates individual word parts (e.g., "pre", "eclampsia")
        }
      }
    }
  },
  "mappings": {
    "properties": {
      "content": {
        "type": "text",
        "analyzer": "hyphen_smart_analyzer",
        "search_analyzer": "hyphen_smart_analyzer" // Use the same logic for indexing and searching
      }
    }
  }
}

Step 2: How It Solves Your Scenarios

Let’s break down the behavior for your key use cases:

Case 1: Hyphenated ↔ Non-Hyphenated Terms

When indexing pre-eclampsia, the analyzer generates these tokens:

  • pre-eclampsia (original term, preserved)
  • pre (split word part)
  • eclampsia (split word part)
  • preeclampsia (combined word parts)

When indexing preeclampsia, the standard tokenizer generates just preeclampsia.

Now, searching for either term will match both documents:

  • Searching pre-eclampsia includes the preeclampsia token in its query, matching the non-hyphenated document.
  • Searching preeclampsia matches the combined token from the hyphenated document.

Case 2: Space-Searched Phrases ↔ Hyphenated Terms

When indexing jean-claude, the analyzer generates:

  • jean-claude (original)
  • jean (split part)
  • claude (split part)
  • jeanclaude (combined)

When searching jean claude, the query splits into jean and claude tokens—both of which exist in the indexed tokens for jean-claude, so the document matches.

Conversely, searching jean-claude will also match documents with jean claude, since the query generates jean and claude tokens that align with the space-separated text.

Step 3: Verify the Analyzer

You can test the analyzer’s output directly with the _analyze endpoint:

POST /hyphenated_docs/_analyze
{
  "analyzer": "hyphen_smart_analyzer",
  "text": "happy-go-lucky"
}

This returns tokens: happy-go-lucky, happy, go, lucky, happygolucky—covering all possible matching scenarios for that term.

Key Advantages Over Application-Layer Workarounds

  • No manual synonym lists: Works for all hyphenated terms automatically, even with large, diverse datasets.
  • Preserves phrase accuracy: The word_delimiter_graph filter maintains term position information, so phrase queries (e.g., "happy go lucky") still work correctly.
  • Single source of truth: All logic lives in Elasticsearch, avoiding inconsistencies between application code and search behavior.

内容的提问来源于stack exchange,提问作者michael

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:27:41