Elasticsearch Python客户端停用词与过滤器配置失效问题排查
问题描述
我正在用Python客户端学习Elasticsearch,已成功创建索引并实现查询功能。但遇到问题:尽管在配置中使用了stemmer和stop words,用纯停用词文本测试查询时,仍返回了本应无结果的匹配,请问哪里配置错误?
配置与代码
# 创建自定义词干过滤器 STEMMER_FILTER = { "type":"stemmer", "language": "english", } # 创建自定义停用词过滤器 STOPWORD_FILTER = { "type":"stop", "stopwords":"_english_", "ignore_case": True, } # 创建同义词过滤器 SYNONYM_FILTER = { "type":"synonym", "synonyms":[ "i-pod, ipod", "universe, cosmos" ] } custom_synonyms = { "type": "synonym_graph", "synonyms": [ "mind, brain", "brain storm, brainstorm, envisage" ] } custom_index_analyzer = { "tokenizer": "whitespace", #"standard", "filter": [ "lowercase", "asciifolding", "stemmer", # 如何使用自定义词干过滤器? "stop"]} # 如何使用自定义停用词过滤器? custom_search_analyzer = { "tokenizer": "whitespace", #"standard" "filter": [ "lowercase", "asciifolding", "stemmer", "stop"]} INDEX_BODY = { "settings": { "index": { "analysis": { "analyzer": { "custom_index_time_analyzer": custom_index_analyzer, "custom_search_time_analyzer": custom_search_analyzer }, "filter": {"my_graph_synonyms": custom_synonyms, "english_stemmer": STEMMER_FILTER, "english_stop": STOPWORD_FILTER, "synonym": SYNONYM_FILTER}, } } }, "mappings": { "properties": { "que_op": {"type": "text", "analyzer": "custom_index_time_analyzer", "search_analyzer": "custom_search_time_analyzer"} }}} es.indices.create(index="questions", mappings=INDEX_BODY["mappings"], settings=INDEX_BODY["settings"]) bulk_data = [] for i,row in tqdm(df.iterrows()): bulk_data.append( { "_index": index_name, "_id": i, "_source": { "que_op": row["que_op"] } } ) bulk(es, bulk_data)
查询代码
def full_text_search(index_name:str, query_string:str, search_on_field:str = 'que_op', size:int = 10): query = {"match": {search_on_field: query_string}} return es.search(index = index_name, query = query, size = size, pretty = True) full_text_search("questions", "a an and are as at be but by for if in into is it no not of on or such that the their then there these they this to was will with", size = 3)
问题原因与解决方法
核心错误
你在自定义分析器中引用的是ES默认的stemmer和stop过滤器,但你自己配置的自定义过滤器名字是english_stemmer和english_stop——分析器根本没用到你定义的停用词和词干规则,导致停用词过滤完全不生效。
ES默认的stop过滤器默认没有启用任何停用词列表,所以你输入的纯停用词会被完整保留并参与匹配,自然会返回包含这些词的文档。
修正步骤
- 修改分析器配置,将filter列表中的
stemmer和stop替换为你自定义的过滤器名称english_stemmer和english_stop:
custom_index_analyzer = { "tokenizer": "whitespace", "filter": [ "lowercase", "asciifolding", "english_stemmer", # 引用自定义词干过滤器 "english_stop" # 引用自定义停用词过滤器 ] } custom_search_analyzer = { "tokenizer": "whitespace", "filter": [ "lowercase", "asciifolding", "english_stemmer", "english_stop" ] }
- 重新创建索引并导入数据:分析器配置是在索引创建时生效的,已存在的索引无法修改分析器,所以需要删除旧索引,用修正后的配置重新创建,再重新导入批量数据。
完成以上操作后,纯停用词的查询就会被过滤,不会返回任何结果。
内容的提问来源于stack exchange,提问作者Deshwal
相关产品推荐
相关产品推荐

