Elasticsearch分析器在API测试正常但搜索查询无预期结果
Elasticsearch查询时无法返回分析器分词结果的问题解决
问题背景
我在索引的settings和mapping中配置了自定义分析器regex_analyzer,配置内容如下:
{ "settings": { "index": { "analysis": { "analyzer": { "synonym_analyzer": { "tokenizer": "standard", "filter": [ "lowercase" ] }, "regex_analyzer": { "tokenizer": "regex_tokenizer", "filter": [ "lowercase" ] } }, "tokenizer": { "regex_tokenizer": { "type": "pattern", "pattern": "((\\b|\\s|\\.|,)[a-z](\\b|\\s |\\.|,)){3,}", "group": 0 } } } } }, "mappings": { "properties": { "transcript_data": { "properties": { "transcript": { "type": "text", "fields": { "keyword": { "type": "keyword" }, "regex": { "type": "text", "analyzer": "regex_analyzer", "search_analyzer": "regex_analyzer" } } } } } } } }
直接调用_analyze API测试该分析器时,能得到符合预期的分词结果:
POST myIndex/_analyze { "analyzer": "regex_analyzer", "text": " this article is talking about l a z r and b k k t ...." }
响应结果:
{ "tokens" : [ { "token" : " b k k t", "start_offset" : 7971, "end_offset" : 7979, "type" : "word", "position" : 0 }, { "token" : " l a z r", "start_offset" : 8350, "end_offset" : 8358, "type" : "word", "position" : 1 } ] }
但使用match_all查询并指定返回transcript_data.transcript.regex字段时,返回的是完整的原始文本,而非分词后的数组:
GET myIndex/_search { "query": { "match_all": { } }, "fields": [ "transcript_data.transcript.regex" ] }
响应结果中fields部分:
"fields" : { "transcript_data.transcript.regex" : [ " this article is talking about l a z r and b k k t ...." ] }
我原本期望这个字段返回与_analyze一致的分词结果。
原因分析
Elasticsearch的fields参数返回的是字段的原始值(仅经过索引时的字符过滤器处理,不会拆解成分词),倒排索引中的分词tokens是用于搜索匹配的内部数据,默认不会通过_search的fields直接返回。
解决方案
方案1:使用Term Vectors API获取分词结果
Term Vectors API可以直接返回指定文档中某个字段的分词信息,包括token、位置、偏移量等,完全匹配_analyze的输出格式。调用方式如下:
GET myIndex/_termvectors/46?fields=transcript_data.transcript.regex
其中46是目标文档的_id,响应结果会包含该字段经过regex_analyzer处理后的所有分词tokens。
方案2:在索引时存储分词结果(可选)
如果需要在搜索时直接返回分词结果,可以通过Ingest Pipeline在文档写入前完成分词,将结果存入一个新字段:
- 创建分词处理的pipeline:
PUT _ingest/pipeline/extract_regex_tokens { "processors": [ { "script": { "source": """ def analyzer = ctx._index + '_' + ctx._type + '#regex_analyzer'; def tokens = analyze(analyzer, ctx.transcript_data.transcript); ctx.transcript_data.regex_tokens = tokens.stream().map(token -> token.token).collect(Collectors.toList()); """ } } ] }
- 写入文档时指定pipeline:
PUT myIndex/_doc/46?pipeline=extract_regex_tokens { "doc_type" : "post", "transcript_data" : { "transcript" : "this article is talking about l a z r and b k k t ...." }, "join_field" : { "name" : "video", "parent" : "anonymouse" } }
之后搜索时就可以直接返回transcript_data.regex_tokens字段,得到分词后的数组。
内容的提问来源于stack exchange,提问作者Exorcismus
相关产品推荐
相关产品推荐

