You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Elasticsearch如何配置多设置实现停用词过滤及同义词检索

Elasticsearch自定义分词配置实现停用词过滤+同义词匹配方案

问题根因

  • 现有players索引未针对name、description字段绑定加载了停用词、同义词规则的自定义分析器,索引写入和检索阶段均未生效对应规则
  • 默认分词器无法识别synonym.txt中football、soccer的同义词映射,也不会过滤stopwords.txt中定义的停用词"was",最终导致误召回id=3文档、漏召回id=2文档

前置配置准备

  • 将stopwords.txt、synonym.txt上传至Elasticsearch集群所有节点的config目录下,确保ES运行账号有文件读取权限
  • 确认synonym.txt内容配置正确:单行写入football,soccer
  • 确认stopwords.txt中已逐行录入所有停用词,包含本次测试用到的was

第一步:创建索引并配置自定义分析器、字段映射

注意:如果已经创建过同名players索引,需要先执行DELETE /players删除旧索引后再执行创建操作,已存在的字段映射不支持直接修改分词器配置。

PUT /players
{
  "settings": {
    "analysis": {
      "analyzer": {
        "custom_text_analyzer": {
          "type": "custom",
          "tokenizer": "standard",
          "filter": [
            "lowercase",
            "my_stop_filter",
            "my_synonym_filter"
          ]
        }
      },
      "filter": {
        "my_stop_filter": {
          "type": "stop",
          "stopwords_path": "stopwords.txt",
          "ignore_case": true
        },
        "my_synonym_filter": {
          "type": "synonym",
          "synonyms_path": "synonym.txt"
        }
      }
    }
  },
  "mappings": {
    "properties": {
      "id": {"type": "integer"},
      "name": {
        "type": "text",
        "analyzer": "custom_text_analyzer",
        "search_analyzer": "custom_text_analyzer"
      },
      "description": {
        "type": "text",
        "analyzer": "custom_text_analyzer",
        "search_analyzer": "custom_text_analyzer"
      },
      "type": {"type": "keyword"}
    }
  }
}

配置说明:

  • 自定义分析器按顺序执行标准分词、小写转换、停用词过滤、同义词扩展流程
  • 停用词过滤器直接加载本地stopwords.txt词表,命中停用词的词元会被直接丢弃,既不会写入倒排索引,也不会参与检索匹配
  • 同义词过滤器加载本地synonym.txt规则,football和soccer会在索引和检索阶段做双向扩展,搜索任意一个词都能匹配到两个词对应的文档

第二步:写入测试数据

from elasticsearch import Elasticsearch
es = Elasticsearch("http://localhost:9200")

abc = [
{'id':1, 'name': 'christiano ronaldo', 'description': 'football@fifa.com', 'type': 'football'},
{'id':2, 'name': 'lionel messi', 'description': 'soccer@fifa.com','type': 'soccer'},
{'id':3, 'name': 'sachin', 'description': 'was', 'type': 'cricket'}
]

for doc in abc:
    es.index(index="players", document=doc)

# 刷新索引使写入立即生效
es.indices.refresh(index="players")

第三步:执行检索验证

无需修改原有检索逻辑,直接执行查询即可:

resp = es.search(index="players",body={
"query": {
"query_string": {
"fields": ["name^2","description^2"],
"query": "was football*"
}
}})
print(resp)

效果说明

  • 检索词中的停用词was会被分析器直接过滤,不参与匹配,因此description字段值为was的id=3文档不会被召回
  • 检索词football会触发同义词扩展规则,自动匹配soccer对应的文档,因此id=2的文档会被正常召回
  • 最终返回id=1、id=2两条文档,和预期结果完全一致

内容的提问来源于stack exchange,提问作者sim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 21:39:03