You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Solr 9查询DenseVectorField时获取距离分数?

在Solr 9.3.0中获取KNN向量查询的每个文档点积距离

我在Solr 9.3.0中创建了诗歌与童谣索引,使用DenseVectorField配置了基于dot_product的KNN向量字段,查询时能匹配到相关文档,但无法获取每个匹配文档对应的点积距离——目前仅能在debug响应中看到一个整体的数值,没有每个文档的单独距离数据。

我的字段配置如下:

<fieldType name="knn_vector" class="solr.DenseVectorField" vectorDimension="384"
  similarityFunction="dot_product"  knnAlgorithm="hnsw" 
  hnswMaxConnections="16" hnswBeamWidth="50"/>
<field name="bge_small_vector" type="knn_vector" indexed="true" stored="true"/>

查询的Python代码:

import pysolr
from sentence_transformers import SentenceTransformer
import pprint

pp = pprint.PrettyPrinter(indent=4, width=100)
solr = pysolr.Solr('http://localhost:8983/solr/docindex')
model = SentenceTransformer('BAAI/bge-small-en-v1.5')
document = '''Three blind mice. Three blind mice.
See how they run. See how they run.
They all ran after the farmer's wife,
Who cut off their tails with a carving knife.
Did you ever see such a sight in your life
As three blind mice?'''
embedding = model.encode(document, normalize_embeddings=True, convert_to_numpy=True)

solr_response=solr.search(
    q=r'{!knn f=bge_small_vector topK=10}[' + ",".join([f'{a:.12f}' for a in embedding]) + ']',
    rows=10,
    start=0,
    debugQuery="true",
    wt='json')

for item in solr_response:
   pp.pprint(item)

目前debug响应中仅能看到全局数值,没有每个文档的距离:

{
    'QParser': 'KnnQParser',
    'explain': {'': '\n**0.81944466 = within top 10**\n'},
    'parsedquery': 'KnnVectorQuery(KnnVectorQuery:bge_small_vector[-0.02721269,...][10])',
    'parsedquery_toString': 'KnnVectorQuery:bge_small_vector[-0.02721269,...][10]',
    ...
}

解决方案

方法1:通过fl参数返回score字段

当向量字段设置similarityFunction="dot_product"时,Solr会将点积值直接作为文档的score返回。只需在查询中添加fl=*,score参数,就能在每个匹配文档中获取对应的点积距离。

修改后的Python查询代码:

solr_response=solr.search(
    q=r'{!knn f=bge_small_vector topK=10}[' + ",".join([f'{a:.12f}' for a in embedding]) + ']',
    rows=10,
    start=0,
    fl='*,score',  # 添加此参数返回点积对应的score
    wt='json')

for item in solr_response:
    # 每个item中会包含'score'字段,值为该文档与查询向量的点积
    pp.pprint(f"文档ID: {item.get('id')}, 点积距离: {item.get('score')}")

方法2:使用函数查询自定义字段名(可选)

如果需要更明确的字段名(比如dot_product_score),可以通过Solr的函数查询直接计算点积并返回:

vector_str = ",".join([f'{a:.12f}' for a in embedding])

solr_response=solr.search(
    q=r'{!knn f=bge_small_vector topK=10}[' + vector_str + ']',
    rows=10,
    start=0,
    # 用func语法计算点积并指定别名
    fl='*,{!func}dot_product(bge_small_vector, [' + vector_str + ']) as dot_product_score',
    wt='json')

这样每个文档会包含dot_product_score字段,存储对应的点积值。


注意事项

  • 由于编码时设置了normalize_embeddings=True,此时点积值等价于余弦相似度,范围在[-1, 1]之间,值越大表示文档与查询向量越相似。
  • Solr的KNN查询默认不会在debug的explain中输出每个文档的详细距离,因此通过fl参数获取score是最直接的方式。

内容的提问来源于stack exchange,提问作者Peter

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 22:05:00