如何通过Elasticsearch内部文档ID定位触发报错的目标文档?
问题场景
在Rails应用中基于Elasticsearch索引执行全文档文本搜索时,触发如下报错:
The length [3618270] of field [text] in doc[126737]/index[my-index] exceeds the [index.highlight.max_analyzed_offset] limit [1000000]. To avoid this error, set the query parameter [max_analyzed_offset] to a value less than index setting [1000000] and this will tolerate long field values by truncating them.
报错里的doc[126737]是Elasticsearch底层Lucene的内部文档ID,并非项目自定义的文档ID,直接通过http://localhost:9200/my-index/_doc/126737查询会返回未找到,以下是具体定位方法:
定位步骤
1. 通过脚本查询匹配内部ID
向Elasticsearch发送搜索请求,用脚本直接匹配Lucene内部文档ID,获取对应文档的自定义ID及完整内容:
请求地址:http://localhost:9200/my-index/_search
请求方式:POST/PUT
请求体:
{ "query": { "bool": { "filter": { "script": { "script": "doc._id == 126737" } } } }, "_source": true }
将126737替换为报错中的实际内部ID数值即可。执行后返回的结果里,会包含该文档的自定义_id以及所有字段内容,以此定位到超长text字段的目标文档。
2. (可选)指定分片缩小查询范围
若已知该内部ID所在的分片编号,可在请求地址后添加参数缩小查询范围,比如分片号为0时:http://localhost:9200/my-index/_search?preference=_shards:0
请求体与上述一致,能提升查询效率。
后续处理
定位到目标文档后,在Rails API的摄入逻辑中添加text字段截断处理,示例代码:
# 在数据入库前截断text字段,控制在1000000字符以内 document.text = document.text.truncate(999999, omission: '')
内容的提问来源于stack exchange,提问作者William Dewey

