You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Elasticsearch中拼接词搜索问题的解决方案咨询

问题描述

基于Elasticsearch搭建商品搜索引擎时,遇到拼接词搜索不一致的问题:

  • 搜索「Smart watch」时,可同时匹配标题含「smartwatch」和「smart watch」的商品
  • 搜索「smartwatch」时,仅能匹配标题含「smartwatch」的商品,无法匹配带空格的「smart watch」变体
当前索引配置
config = {
 "settings": {
    "analysis": {
      "analyzer": {
        "nGram_analyzer": {
          "type": "custom",
          "tokenizer": "whitespace",
          "char_filter":["html_strip","custom_char_filter","space_maker_2", "space_maker_3" ],
          "filter": [
            "lowercase",
            "asciifolding",
            "nGram_filter"
          ]
        },
        "whitespace_analyzer": {
          "type": "custom",
          "tokenizer": "whitespace",
          "char_filter": ["space_maker_2", "space_maker_3"
          ],
          "filter": [
            "lowercase",
            "asciifolding",
            "synonym_apply",
            "special_stopwards"
          ]
        }
      },
      "char_filter": {
        "custom_char_filter": {
          "type": "mapping",
          "mappings": [
            "$ => dollar"
          ]
        },
        "space_maker_1": {
          "type": "pattern_replace", 
          "pattern": "(?<=[a-z])(?=[A-Z])|(?<=[A-Z])(?=[a-z])",
          "replacement": " "
        },
        "space_maker_2": {
          "type": "pattern_replace",
          "pattern": "(?<=\\p{Digit})(?=\\p{Alpha})|(?<=\\p{Alpha})(?=\\p{Digit})",
          "replacement": " "
        },
        "space_maker_3": {
          "type": "pattern_replace",
          "pattern": "(?<=[a-zA-Z0-9])(?=[^a-zA-Z0-9])|(?<=[^a-zA-Z0-9])(?=[a-zA-Z0-9])",
          "replacement": " "
        }
      },
       "filter": {
        "nGram_filter": {
          "type": "edge_ngram",
          "min_gram": 2,
          "max_gram": 20,
          "token_chars": [
            "letter",
            "digit",
            "punctuation",
            "symbol"
          ]
        },
        "synonym_apply": {
            "type": "synonym",
            "lenient": "true",
            "synonyms": [ "kilo, kilogram => kg",
            "buck, dollar => usd"
            ]
          },
        "special_stopwards": {
            "type": "stop",
            "stopwords": [ "ass", "butt" ]
          }
      }
    }
  },

    "mappings": {
        "properties": {
            "brand": {
                "type": "keyword"
            },
            "category": {
                "type": "keyword"
            },
            "tags": {
                "type": "keyword"
            },
            "domain": {
                "type": "keyword"
            },
            "image": {
                "type": "text"
            },
            "purchases": {
                "type": "double"
            },
            "views": {
                "type": "double"
            },
            "price": {
                "type": "double"
            },
            "product_id": {
                "type": "text"
            },
            "product_url": {
                "type": "text"
            },
            "title": {
                "type": "text",
                "analyzer": "nGram_analyzer",
                "search_analyzer": "nGram_analyzer",
            },
            "description": {
                "type": "text"
            },
            "country": {
                "type": "integer"
            },
            "last_seen_date": {
                "type": "text"
            }
        }
    }
}
解决方案

方法1:用word_delimiter_graph过滤器自动处理拼接词

这是通用且高效的方案,通过过滤器将拼接词拆分为子词并保留原词,让两种形式的词在索引时生成重叠的token,实现互相匹配。

  1. 在filter节点新增word_delimiter配置:
"filter": {
  # 保留原有filter(nGram_filter、synonym_apply等)
  "word_delimiter": {
    "type": "word_delimiter_graph",
    "split_on_case_change": true,  # 按大小写拆分驼峰词(如SmartWatch→Smart Watch)
    "generate_word_parts": true,   # 生成子词片段
    "generate_number_parts": true,
    "catenate_words": true,        # 将子词拼接回原词
    "catenate_numbers": true,
    "catenate_all": true,
    "preserve_original": true      # 保留原始token
  }
}
  1. 修改nGram_analyzer:加入space_maker_1处理驼峰词,同时将word_delimiter添加到过滤器链中:
"nGram_analyzer": {
  "type": "custom",
  "tokenizer": "whitespace",
  "char_filter": ["html_strip","custom_char_filter","space_maker_1", "space_maker_2", "space_maker_3"],
  "filter": [
    "lowercase",
    "asciifolding",
    "word_delimiter",
    "nGram_filter"
  ]
}
  1. 重新创建索引并导入数据(分析器修改后必须重新索引才能生效)

方法2:同义词过滤器(适合固定词汇场景)

如果仅需处理特定词对,可维护同义词表实现双向匹配:

  1. 更新synonym_apply的同义词列表:
"synonym_apply": {
  "type": "synonym",
  "lenient": "true",
  "synonyms": [ 
    "kilo, kilogram => kg",
    "buck, dollar => usd",
    "smartwatch, smart watch"  # 添加双向同义词映射
  ]
}
  1. 将synonym_apply加入nGram_analyzer的过滤器列表
    缺点:需手动维护所有词汇对,不适合通用商品搜索场景

方法3:调整查询逻辑(无需修改索引)

若无法重新索引,可在查询时同时搜索两种词形:

{
  "query": {
    "bool": {
      "should": [
        {"match": {"title": "smartwatch"}},
        {"match": {"title": "smart watch"}}
      ]
    }
  }
}

也可在搜索前对查询词预处理,自动生成带空格的变体后再查询。此方法为临时解决方案,长期效率不如修改索引分析器。


内容的提问来源于stack exchange,提问作者RefiPeretz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 14:59:51