Elasticsearch TF-IDF计算异常与分词权重优化技术咨询
背景信息
我有一个存储旅游信息的Elasticsearch索引,可搜索字段会被复制到searchable_keys字段中,当前仅包含name字段。索引定义如下:
{ "settings":{ "analysis":{ "analyzer":{ "my_analyzer":{ "filter":[ "lowercase" ], "type":"custom", "tokenizer":"my_tokenizer" } }, "tokenizer":{ "my_tokenizer":{ "token_chars":[ "letter", "digit" ], "type":"edge_ngram", "min_gram":3, "max_gram":20 } } } }, "mappings":{ "properties":{ "entry_id":{ "type":"keyword" }, "workspace_id":{ "type":"keyword" }, "name":{ "type":"text", "copy_to":"searchable_keys" }, "searchable_keys":{ "type":"text", "analyzer":"my_analyzer" } } } }
执行以下查询:
{ "explain":true, "query":{ "match":{ "searchable_keys":{ "query":"dog", "operator":"AND" } } } }
得到异常结果:名称为• Private Emerald Lake & Dogsledding Tour •的文档得分为3.7377324,名称为Skagway Sled Dog and Musher's Camp的文档得分为3.718998,详细返回结果如下:
[ { "_index":"tours", "_id":"018bb59a-bc8c-76a2-9e76-eaf747bac7c1", "_score":3.7377324, "_source":{ "entry_id":"018bb59a-bc8c-76a2-9e76-eaf747bac7c1", "workspace_id":"018bb598-708a-7e8d-8995-b30cf0aba239", "name":"• Private Emerald Lake & Dogsledding Tour •", "type":"Tour" }, "_explanation":{ "value":3.7377324, "description":"weight(searchable_keys:dog in 68) [PerFieldSimilarity], result of:", "details":[ { "value":3.7377324, "description":"score(freq=1.0), computed as boost * idf * tf from:", "details":[ { "value":2.2, "description":"boost", "details":[ ] }, { "value":4.017076, "description":"idf, computed as log(1 + (N - n + 0.5) / (n + 0.5)) from:", "details":[ { "value":6, "description":"n, number of documents containing term", "details":[ ] }, { "value":360, "description":"N, total number of documents with field", "details":[ ] } ] }, { "value":0.4229368, "description":"tf, computed as freq / (freq + k1 * (1 - b + b * dl / avgdl)) from:", "details":[ { "value":1.0, "description":"freq, occurrences of term within document", "details":[ ] }, { "value":1.2, "description":"k1, term saturation parameter", "details":[ ] }, { "value":0.75, "description":"b, length normalization parameter", "details":[ ] }, { "value":23.0, "description":"dl, length of field", "details":[ ] }, { "value":19.447222, "description":"avgdl, average length of field", "details":[ ] } ] } ] } ] } }, { "_index":"tours", "_id":"018bb598-e50e-7d6d-a639-97ed40bb2ee7", "_score":3.718998, "_source":{ "entry_id":"018bb598-e50e-7d6d-a639-97ed40bb2ee7", "workspace_id":"018bb598-708a-7e8d-8995-b30cf0aba239", "name":"Skagway Sled Dog and Musher's Camp", "type":"Tour" }, "_explanation":{ "value":3.718998, "description":"weight(searchable_keys:dog in 105) [PerFieldSimilarity], result of:", "details":[ { "value":3.718998, "description":"score(freq=1.0), computed as boost * idf * tf from:", "details":[ { "value":2.2, "description":"boost", "details":[ ] }, { "value":3.3953834, "description":"idf, computed as log(1 + (N - n + 0.5) / (n + 0.5)) from:", "details":[ { "value":11, "description":"n, number of documents containing term", "details":[ ] }, { "value":342, "description":"N, total number of documents with field", "details":[ ] } ] }, { "value":0.49786824, "description":"tf, computed as freq / (freq + k1 * (1 - b + b * dl / avgdl)) from:", "details":[ { "value":1.0, "description":"freq, occurrences of term within document", "details":[ ] }, { "value":1.2, "description":"k1, term saturation parameter", "details":[ ] }, { "value":0.75, "description":"b, length normalization parameter", "details":[ ] }, { "value":15.0, "description":"dl, length of field", "details":[ ] }, { "value":19.052631, "description":"avgdl, average length of field", "details":[ ] } ] } ] } ] } } ]
技术疑问
- 为何两个文档的idf值不同?按理论同一词汇的idf在整个文档集合中应一致,我的理解是否有误?
- 当前tf使用的公式为何与常规的「词频/文档总词数」不同?
- 如何实现包含独立词"dog"的文档,比分词中含"dog"子串的文档得分更高,同时保留edge n-gram分词带来的子串搜索能力?
问题解答
1. 为何两个文档的idf值不同?
你的理解没错,同一词汇的idf理论上应该一致,但问题出在分片级别的idf计算。Elasticsearch默认会在每个分片上独立计算idf值,而非整个索引全局计算。如果两个文档位于不同分片,每个分片的文档总数(N)和包含目标词的文档数(n)存在差异,就会导致idf结果不同。
从你的结果能看到,第一个文档所在分片的N是360、n是6;第二个文档所在分片的N是342、n是11,显然是两个分片的统计数据,所以idf计算结果出现了差异。
2. 当前tf使用的公式为何与常规的「词频/文档总词数」不同?
你看到的是BM25相似度算法的tf计算公式,Elasticsearch从7.0版本开始,默认相似度算法已经从TF-IDF换成了BM25。BM25的tf公式引入了饱和度参数k1和长度归一化参数b,目的是避免TF-IDF中词频过高导致得分过度膨胀的问题,同时兼顾文档长度对得分的影响。
公式freq / (freq + k1 * (1 - b + b * dl / avgdl))中:
freq是词在文档中的出现次数k1控制词频的饱和程度(默认1.2),值越大,词频对得分的影响越接近TF-IDFb控制文档长度对得分的影响(默认0.75),值为0时忽略文档长度,值为1时完全基于文档长度归一化dl是当前文档的字段长度,avgdl是索引中该字段的平均长度
3. 如何实现独立词"dog"的文档得分更高,同时保留edge n-gram的子串搜索能力?
可以通过多字段映射+布尔查询加权的方式实现:
- 第一步:修改
searchable_keys的映射,新增一个使用标准分词器的子字段(比如searchable_keys.exact),用来匹配完整的独立词;同时保留原有的edge n-gram分词字段用于子串搜索。
修改后的映射片段:"searchable_keys":{ "type":"text", "analyzer":"my_analyzer", "fields":{ "exact":{ "type":"text", "analyzer":"standard" } } } - 第二步:使用
bool查询,同时查询两个字段,给exact字段更高的权重(boost),确保匹配完整词的文档得分更高。
查询示例:{ "query":{ "bool":{ "should":[ { "match":{ "searchable_keys.exact":{ "query":"dog", "boost":3 } } }, { "match":{ "searchable_keys":{ "query":"dog" } } } ] } } }
这样,包含独立"dog"词的文档会同时匹配两个字段,获得更高的综合得分;而仅包含"dog"子串的文档只会匹配edge n-gram字段,得分较低,同时子串搜索的能力也得到了保留。
内容的提问来源于stack exchange,提问作者Prosto_Oleg

