Azure AI Search评分逻辑疑问:相同匹配为何得分不同?
Azure AI Search搜索评分异常原因解析
一、第一次搜索:"motel"的评分差异
搜索代码
results = search_client.search(search_text="motel", select='HotelId,HotelName,Rating', order_by='Rating desc', include_total_count=True) print ('Total Documents Matching Query:', results.get_count()) for result in results: print(result["@search.score"]) print("{}: {} - {} rating".format(result["HotelId"], result["HotelName"], result["Rating"]))
搜索结果
Total Documents Matching Query: 2 0.6099695 2: Twin Dome Motel - 3.6 rating 0.25316024 1: Secret Point Motel - 3.6 rating
原因解析
两个酒店名称都仅含1次"motel",评分差异源于Azure AI Search默认使用的BM25评分算法的字段长度归一化机制:
- BM25会根据字段长度调整词的权重:相同匹配词在更短的字段中,权重占比更高——因为短文本中的词对文档主题的贡献度更大。
- "Twin Dome Motel"的字段字符长度比"Secret Point Motel"更短,因此"motel"在该字段中的权重更高,最终评分也更高。
二、第二次搜索:"what hotel has a good restaurant on site"的评分异常
搜索代码
results = search_client.search(query_type='simple', search_text="what hotel has a good restaurant on site" , select='HotelName,HotelId,Description') for result in results: print(result["@search.score"]) print(result["HotelName"])
搜索结果
2.1393623 Sublime Cliff Hotel 1.9309065 Twin Dome Motel 1.5589908 Triple Landscape Hotel 0.7704947 Secret Point Motel
原因解析
这个结果的异常可以从两个核心点解释:
Simple查询的分词与停用词处理
Simple查询模式会自动过滤无实际语义的常用停用词,你的搜索词中what、has、a、on、site都会被过滤,最终实际参与匹配的词是hotel、good、restaurant:good未在任何文档中出现,但Simple查询是OR逻辑,缺少匹配词不会排除文档,只是该词不贡献评分。hotel在所有4个酒店的文档中都有出现(HotelName或Description字段),因此所有文档都能匹配。
BM25的词频与逆文档频率(IDF)计算
hotel是所有文档都包含的高频词,逆文档频率(IDF)极低,单个匹配的权重不高,但如果某个文档中hotel出现次数更多,累计权重会更高:- Sublime Cliff Hotel的Description中多次出现"Hotel"/"hotel",Twin Dome Motel的Description也有多次提及,因此这两个文档的
hotel累计评分更高。
- Sublime Cliff Hotel的Description中多次出现"Hotel"/"hotel",Twin Dome Motel的Description也有多次提及,因此这两个文档的
restaurant仅在Triple Landscape Hotel中出现,虽然它的IDF很高(仅单个文档包含),但该文档中hotel的出现次数远少于前两个文档,总评分是hotel的累计评分加上restaurant的评分,最终总评分低于前两个文档。
内容的提问来源于stack exchange,提问作者Manu Chadha
相关产品推荐
相关产品推荐

