认知搜索多索引查询分数异常问题及归一化方案咨询
多索引向量搜索分数不一致问题处理
问题背景
我有近100个存储不同来源信息的索引,所有数据统一使用adda2做嵌入。目前通过遍历索引列表逐个执行查询,但得到的跨索引搜索分数不符合预期——本该最相关的CORRECT-idx平均分数反而最低。
当前实现代码
遍历索引查询逻辑
index_client = SearchIndexClient( endpoint=pierre_itfunds_endpoint, credential=pierre_itfunds_credential ) index_client.list_indexes() rows_list = [] for index in indexes: search_client = SearchClient(search_service_endpoint, index, credential) vector_query = VectorizedQuery(vector=search_vector, k_nearest_neighbors=3, fields="content_vector") results = search_client.search( search_text=query, vector_queries= [vector_query], select=["title", "text"], top=3 ) for row in results: dict1 = {} dict1.update({'index':index, 'score':row['@search.score'], 'title':row['title'], 'text':row['text']}) rows_list.append(dict1) res = pd.DataFrame(rows_list)
计算索引平均分数
grouped = res.groupby('index')['score'].agg(['mean']) grouped
不符合预期的结果
| index_name | avg_score |
|---|---|
| worng-idx | 8.3763725667 |
| other1-idx | 4.5701991333 |
| other2-idx | 4.2485168 |
| other3-idx | 3.5756512667 |
| CORRECT-idx | 2.5451367667 |
| ... | ... |
需求
我原本以为使用相同嵌入器后,跨索引的余弦距离会保持一致,虽然清楚余弦距离与搜索分数无关,但还是希望找到可行的多索引搜索方式,或者对分数做归一化处理,让最高分对应最相关的结果。
内容的提问来源于stack exchange,提问作者filippo
相关产品推荐
相关产品推荐

