重复索引同一Elasticsearch文档为何引发统计数与得分异常?
1. 初始操作与查询结果
我们创建索引并写入文档:
POST sample-index-test/_doc/1 { "first_name": "James", "last_name" : "Osaka" }
此时索引中仅存在1个文档,调用_explain API执行match查询:
GET sample-index-test/_explain/1 { "query": { "match": { "first_name": "James" } } }
_explain API返回核心结果:
- score : 0.2876821
- number of documents containing term : 1
- total number of documents with field : 1
完整返回内容:
{ "_index" : "sample-index-test", "_type" : "_doc", "_id" : "1", "matched" : true, "explanation" : { "value" : 0.2876821, "description" : "weight(first_name:james in 0) [PerFieldSimilarity], result of:", "details" : [ { "value" : 0.2876821, "description" : "score(freq=1.0), computed as boost * idf * tf from:", "details" : [ { "value" : 2.2, "description" : "boost", "details" : [ ] }, { "value" : 0.2876821, "description" : "idf, computed as log(1 + (N - n + 0.5) / (n + 0.5)) from:", "details" : [ { "value" : 1, "description" : "n, number of documents containing term", "details" : [ ] }, { "value" : 1, "description" : "N, total number of documents with field", "details" : [ ] } ] }, { "value" : 0.45454544, "description" : "tf, computed as freq / (freq + k1 * (1 - b + b * dl / avgdl)) from:", "details" : [ { "value" : 1.0, "description" : "freq, occurrences of term within document", "details" : [ ] }, { "value" : 1.2, "description" : "k1, term saturation parameter", "details" : [ ] }, { "value" : 0.75, "description" : "b, length normalization parameter", "details" : [ ] }, { "value" : 1.0, "description" : "dl, length of field", "details" : [ ] }, { "value" : 1.0, "description" : "avgdl, average length of field", "details" : [ ] } ] } ] } ] } }
2. 多次更新后的异常结果
随后在数秒内多次执行同一索引请求(更新文档):
POST sample-index-test/_doc/1 { "first_name": "James", "last_name" : "Cena" }
再次调用相同的_explain API后,得到了不同的得分以及异常的文档统计数:
- score : 0.046520013
- number of documents containing term : 10
- total number of documents with field : 10
完整返回内容:
{ "_index" : "sample-index-test", "_type" : "_doc", "_id" : "1", "matched" : true, "explanation" : { "value" : 0.046520013, "description" : "weight(first_name:james in 0) [PerFieldSimilarity], result of:", "details" : [ { "value" : 0.046520013, "description" : "score(freq=1.0), computed as boost * idf * tf from:", "details" : [ { "value" : 2.2, "description" : "boost", "details" : [ ] }, { "value" : 0.046520017, "description" : "idf, computed as log(1 + (N - n + 0.5) / (n + 0.5)) from:", "details" : [ { "value" : 10, "description" : "n, number of documents containing term", "details" : [ ] }, { "value" : 10, "description" : "N, total number of documents with field", "details" : [ ] } ] }, { "value" : 0.45454544, "description" : "tf, computed as freq / (freq + k1 * (1 - b + b * dl / avgdl)) from:", "details" : [ { "value" : 1.0, "description" : "freq, occurrences of term within document", "details" : [ ] }, { "value" : 1.2, "description" : "k1, term saturation parameter", "details" : [ ] }, { "value" : 0.75, "description" : "b, length normalization parameter", "details" : [ ] }, { "value" : 1.0, "description" : "dl, length of field", "details" : [ ] }, { "value" : 1.0, "description" : "avgdl, average length of field", "details" : [ ] } ] } ] } ] } }
问题
为何Elasticsearch在索引仅含单个文档的情况下,会增加“total number of documents with field”和“number of documents containing term”的统计数?
解答
这是因为Elasticsearch的倒排索引统计信息(包括文档数、词频等)并非实时更新,而是基于分片的段(Segment)合并机制维护的。
当你多次执行同一ID的文档更新请求时,Elasticsearch不会直接修改原文档,而是标记原文档为删除状态,再新增一个新版本的文档。短时间内多次更新会生成多个小的段,这些段的统计信息会被暂时独立计算。
_explain API计算IDF(逆文档频率)时,会读取当前所有可见段的统计数据——包括那些已标记为删除但还未被合并清理的旧文档对应的段。所以即使索引中实际只有1个可用文档,统计数会把所有未合并段中的文档(含已删除的旧版本)都算进去,导致n(包含该词的文档数)和N(包含该字段的总文档数)被错误放大。
等段合并操作完成后,已删除的文档会被清理,统计数就会恢复正确数值。你也可以手动调用POST sample-index-test/_forcemerge触发段合并,之后再查询_explain就能看到正常结果。
内容的提问来源于stack exchange,提问作者rajesh

