如何在Solr 9查询DenseVectorField时获取距离分数?
在Solr 9.3.0中获取KNN向量查询的每个文档点积距离
我在Solr 9.3.0中创建了诗歌与童谣索引,使用DenseVectorField配置了基于dot_product的KNN向量字段,查询时能匹配到相关文档,但无法获取每个匹配文档对应的点积距离——目前仅能在debug响应中看到一个整体的数值,没有每个文档的单独距离数据。
我的字段配置如下:
<fieldType name="knn_vector" class="solr.DenseVectorField" vectorDimension="384" similarityFunction="dot_product" knnAlgorithm="hnsw" hnswMaxConnections="16" hnswBeamWidth="50"/> <field name="bge_small_vector" type="knn_vector" indexed="true" stored="true"/>
查询的Python代码:
import pysolr from sentence_transformers import SentenceTransformer import pprint pp = pprint.PrettyPrinter(indent=4, width=100) solr = pysolr.Solr('http://localhost:8983/solr/docindex') model = SentenceTransformer('BAAI/bge-small-en-v1.5') document = '''Three blind mice. Three blind mice. See how they run. See how they run. They all ran after the farmer's wife, Who cut off their tails with a carving knife. Did you ever see such a sight in your life As three blind mice?''' embedding = model.encode(document, normalize_embeddings=True, convert_to_numpy=True) solr_response=solr.search( q=r'{!knn f=bge_small_vector topK=10}[' + ",".join([f'{a:.12f}' for a in embedding]) + ']', rows=10, start=0, debugQuery="true", wt='json') for item in solr_response: pp.pprint(item)
目前debug响应中仅能看到全局数值,没有每个文档的距离:
{ 'QParser': 'KnnQParser', 'explain': {'': '\n**0.81944466 = within top 10**\n'}, 'parsedquery': 'KnnVectorQuery(KnnVectorQuery:bge_small_vector[-0.02721269,...][10])', 'parsedquery_toString': 'KnnVectorQuery:bge_small_vector[-0.02721269,...][10]', ... }
解决方案
方法1:通过fl参数返回score字段
当向量字段设置similarityFunction="dot_product"时,Solr会将点积值直接作为文档的score返回。只需在查询中添加fl=*,score参数,就能在每个匹配文档中获取对应的点积距离。
修改后的Python查询代码:
solr_response=solr.search( q=r'{!knn f=bge_small_vector topK=10}[' + ",".join([f'{a:.12f}' for a in embedding]) + ']', rows=10, start=0, fl='*,score', # 添加此参数返回点积对应的score wt='json') for item in solr_response: # 每个item中会包含'score'字段,值为该文档与查询向量的点积 pp.pprint(f"文档ID: {item.get('id')}, 点积距离: {item.get('score')}")
方法2:使用函数查询自定义字段名(可选)
如果需要更明确的字段名(比如dot_product_score),可以通过Solr的函数查询直接计算点积并返回:
vector_str = ",".join([f'{a:.12f}' for a in embedding]) solr_response=solr.search( q=r'{!knn f=bge_small_vector topK=10}[' + vector_str + ']', rows=10, start=0, # 用func语法计算点积并指定别名 fl='*,{!func}dot_product(bge_small_vector, [' + vector_str + ']) as dot_product_score', wt='json')
这样每个文档会包含dot_product_score字段,存储对应的点积值。
注意事项
- 由于编码时设置了
normalize_embeddings=True,此时点积值等价于余弦相似度,范围在[-1, 1]之间,值越大表示文档与查询向量越相似。 - Solr的KNN查询默认不会在debug的explain中输出每个文档的详细距离,因此通过
fl参数获取score是最直接的方式。
内容的提问来源于stack exchange,提问作者Peter
相关产品推荐
相关产品推荐

