Elastic Search 查询时Analyzer失效问题求助
问题重现
- 创建仅包含查询时分析器的索引:
PUT /local_persons { "settings": { "analysis": { "analyzer": { "person_search_analyzer": { "type": "custom", "char_filter": ["remove_special_chars"], "filter": ["lowercase"], "tokenizer": "whitespace" } }, "char_filter": { "remove_special_chars": { "type": "pattern_replace", "pattern": "[^a-zA-Z0-9]", "replacement": "" } } } } }
- 索引包含特殊字符的文档:
PUT /local_persons/_doc/1 { "id": 1, "firstName": "Re'mo", "lastName": "D'souza", "email": "remo.d@test.com", "dateOfBirth": "1973-01-01", "isActive": 1 }
尝试两种查询方式均返回空结果:
- 方法一:使用query_string指定查询分析器
GET /local_persons/_search { "query": { "bool": { "must": [ { "query_string": { "query": "remo", "fields": ["firstName"], "analyzer": "person_search_analyzer" } }, { "query_string": { "query": "dsouza", "fields": ["lastName"], "analyzer": "person_search_analyzer" } } ] } } }- 方法二:使用match查询指定查询分析器
GET /local_persons/_search { "query": { "bool": { "must": [ { "match": { "firstName": { "query": "remo", "analyzer": "person_search_analyzer" } } }, { "match": { "lastName": { "query": "dsouza", "analyzer": "person_search_analyzer" } } } ] } } }
问题原因
核心是索引阶段与查询阶段的词条不匹配:
- 索引文档时,
firstName、lastName字段用Elasticsearch默认标准分析器,会把Re'mo拆成["re", "mo"],D'souza拆成["d", "souza"],仅保留拆分后的小写形式。 - 查询时,自定义
person_search_analyzer会把remo处理成["remo"],dsouza处理成["dsouza"],和索引中存储的词条完全不匹配,因此返回空结果。
解决方案
需要让索引、查询阶段使用一致的处理逻辑,确保生成的词条可匹配,以下是两种常用方案:
方案一:创建索引时指定字段的索引/查询分析器
在索引映射中,给目标字段指定analyzer(索引时使用)和search_analyzer(查询时使用)为自定义的person_search_analyzer:
PUT /local_persons { "settings": { "analysis": { "analyzer": { "person_search_analyzer": { "type": "custom", "char_filter": ["remove_special_chars"], "filter": ["lowercase"], "tokenizer": "whitespace" } }, "char_filter": { "remove_special_chars": { "type": "pattern_replace", "pattern": "[^a-zA-Z0-9]", "replacement": "" } } } }, "mappings": { "properties": { "firstName": { "type": "text", "analyzer": "person_search_analyzer", "search_analyzer": "person_search_analyzer" }, "lastName": { "type": "text", "analyzer": "person_search_analyzer", "search_analyzer": "person_search_analyzer" }, "id": {"type": "integer"}, "email": {"type": "keyword"}, "dateOfBirth": {"type": "date"}, "isActive": {"type": "integer"} } } }
重新插入文档后执行原查询即可匹配结果。
方案二:使用normalizer处理(适合无需分词的场景)
如果字段不需要按空格分词,仅需去除特殊字符并转小写,可将字段设为keyword类型,用normalizer替代分析器:
PUT /local_persons { "settings": { "analysis": { "normalizer": { "person_normalizer": { "type": "custom", "char_filter": ["remove_special_chars"], "filter": ["lowercase"] } }, "char_filter": { "remove_special_chars": { "type": "pattern_replace", "pattern": "[^a-zA-Z0-9]", "replacement": "" } } } }, "mappings": { "properties": { "firstName": { "type": "keyword", "normalizer": "person_normalizer" }, "lastName": { "type": "keyword", "normalizer": "person_normalizer" }, "id": {"type": "integer"}, "email": {"type": "keyword"}, "dateOfBirth": {"type": "date"}, "isActive": {"type": "integer"} } } }
这种方式下,Re'mo会被处理为remo,D'souza处理为dsouza,查询时可直接匹配完整词条。
验证方式
用_analyzeAPI验证分析器处理结果,确保索引与查询的词条一致:
POST /local_persons/_analyze { "analyzer": "person_search_analyzer", "text": "Re'mo" }
预期返回:
{ "tokens": [ { "token": "remo", "start_offset": 0, "end_offset": 5, "type": "word", "position": 0 } ] }
内容的提问来源于stack exchange,提问作者Ajay
相关产品推荐
相关产品推荐

