Elasticsearch自定义分析器停用词过滤与文本存储异常问题
Elasticsearch自定义分析器疑问解答
操作过程
创建索引请求
PUT my_index { "settings": { "index.number_of_shards": 1, "index.number_of_replicas": 0, "analysis": { "analyzer": { "my_custom_analyzer": { "type": "custom", "tokenizer": "standard", "max_token_length": 3, "filter": [ "lowercase", "english_stop", "english_stemmer", "asciifolding" ] } }, "filter": { "english_stemmer": { "type": "stemmer", "stopwords": "english" }, "english_stop": { "type": "stop", "stopwords": "_english_" } } } }, "mappings": { "properties": { "Text": { "type": "text", "analyzer": "my_custom_analyzer" } } } }
添加文档请求
POST my_index/_doc/1 { "Text": "a quick fox jumps over the lazy dog" }
查询请求
GET my_index/_search { "query": { "match": { "Text": { "query": "a the", "operator": "and" } } } }
查询响应
{ "took" : 1, "timed_out" : false, "_shards" : { "total" : 1, "successful" : 1, "skipped" : 0, "failed" : 0 }, "hits" : { "total" : { "value" : 1, "relation" : "eq" }, "max_score" : 0.5753642, "hits" : [ { "_index" : "my_index", "_type" : "_doc", "_id" : "1", "_score" : 0.5753642, "_source" : { "Text" : "a quick fox jumps over the lazy dog" } } ] } }
疑问与解答
疑问1:停用词查询为何仍能匹配结果?
原因是查询阶段会使用字段指定的分析器处理查询文本。你定义的my_custom_analyzer包含停用词过滤器english_stop,当查询文本"a the"被分析时,这两个停用词会被直接过滤掉,最终查询没有有效匹配词。此时match查询(operator: and)会退化为匹配所有文档,所以你能得到结果。
可以用_analyze接口验证分析器对查询文本的处理结果:
GET _analyze { "analyzer": "my_custom_analyzer", "text": "a the" }
返回结果会显示没有生成任何词项,这直接解释了查询的行为。
疑问2:Text字段为何仍存储原始文本?
Elasticsearch中text字段的_source是原始文档的备份,用于查询时返回给用户查看,它不会被分析器修改。真正用于检索的是索引中存储的经过分析后的词项(也就是你预期的去除停用词后的精简版本),这些词项仅用于内部匹配计算,不会直接返回。
可以用_termvectors接口确认索引中的词项:
GET my_index/_doc/1/_termvectors { "fields": ["Text"] }
返回结果里不会包含"a"和"the"这些停用词,证明分析器确实在索引阶段过滤了它们。
内容的提问来源于stack exchange,提问作者No1Lives4Ever
相关产品推荐
相关产品推荐

