You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Weaviate with_near_vector()未返回完全匹配向量记录的问题

问题描述

我有一个存储70万条记录的Weaviate向量类,使用自定义向量。遇到以下问题:

  • 用with_near_vector()传入完全匹配的向量查询时,预期Top结果是距离接近0的匹配记录(node_type为"type2"),但实际返回的最接近记录距离约0.1且node_type是"type1"。
  • 执行带distance:1.0的查询时,仅返回11100条结果(理论上应返回所有符合距离的记录,但受计算限制):
# NOTE: mean_emb是与MyClass中某条记录匹配的numpy数组。
# 理论上distance设为1.0应返回所有70万条记录的距离,但受计算限制无法实现
result = (client.query.get("MyClass", ["message", "node_type", "my_id", "timestamp"])
          .with_near_vector({"vector": mean_emb.tolist(), "distance": 1.0})
          .with_additional(["vector", "distance"]).do())
result = result["data"]["Get"]["MyClass"]
print(len(result))  # 仅返回11,100条距离结果
  • 分页问题:with_offset()偏移量大于100000时失效;with_after()不支持with_near_vector()查询;with_offset()+with_limit()速度极慢。
  • 但通过where过滤node_type为"type2"时,能查到距离约-1.9073486e-06的匹配记录:
where_filter = {"path": ["node_type"], "operator": "Equal", "valueText": "type2"}
result = (client.query.get("MyClass", ["message", "node_type", "my_id", "timestamp"])
          .with_near_vector({"vector": mean_emb.tolist()})
          .with_additional(["distance", "id"]).with_where(where_filter).do())
print(result)

返回结果:

{'data': {'Get': {'MyClass': [{'_additional': {'distance': -1.9073486e-06,
      'id': 'fdb00f95-2c07-462c-84cd-9380c6777801'},
     'my_id': 'Record that matches the vector passed',
     'message': None,
     'node_type': 'type2',
     'timestamp': None},
    {'_additional': {'distance': 0.6122676,
      'id': '0deb152a-eef0-485c-ad6e-c9e29f9a3915'},
     'my_id': 'Another type2 record that doesn't match vector passed',
     'message': None,
     'node_type': 'type2',
     'timestamp': None}]}}}
解决方案

1. 提升ANN搜索精度

Weaviate默认用近似最近邻(ANN)搜索,可能漏掉精确匹配的记录,可调整参数提升召回率:

  • 建索引阶段:修改类的efConstruction参数(值越高索引精度越高,建索引越慢),比如设为512。
  • 查询阶段:在with_near_vector()中通过search_parameters传入ef参数,增大搜索候选集:
result = (client.query.get("MyClass", ["message", "node_type", "my_id", "timestamp"])
          .with_near_vector({
              "vector": mean_emb.tolist(),
              "search_parameters": {"ef": 1000}
          })
          .with_additional(["distance", "id"])
          .do())

2. 结合类型过滤缩小搜索范围

既然明确匹配记录的node_type是"type2",先过滤该类型再做向量查询,既能确保召回目标记录,又能减少查询数据量:

where_filter = {"path": ["node_type"], "operator": "Equal", "valueText": "type2"}
result = (client.query.get("MyClass", ["message", "node_type", "my_id", "timestamp"])
          .with_near_vector({"vector": mean_emb.tolist()})
          .with_where(where_filter)
          .with_additional(["distance", "id"])
          .with_limit(10)
          .do())

3. 分页问题替代方案

对于大数据量向量查询分页,推荐用批次+距离阈值的方式:

all_results = []
current_max_distance = 1.0
batch_size = 1000

while True:
    result = (client.query.get("MyClass", ["message", "node_type", "my_id", "timestamp"])
              .with_near_vector({
                  "vector": mean_emb.tolist(),
                  "distance": current_max_distance
              })
              .with_additional(["distance"])
              .with_limit(batch_size)
              .do())
    hits = result["data"]["Get"]["MyClass"]
    if not hits:
        break
    all_results.extend(hits)
    # 更新下一次查询的最大距离阈值,避免重复获取
    current_max_distance = min(hit["_additional"]["distance"] for hit in hits) - 0.001

4. 验证向量一致性

确认存入的自定义向量与查询用的mean_emb完全一致:

  • 取出匹配记录的向量(通过with_additional(["vector"])),与mean_emb做数值对比,检查是否存在numpy转list时的精度损失
  • 确保存入和查询时向量的维度、数据类型完全匹配

内容的提问来源于stack exchange,提问作者lrthistlethwaite

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 17:33:08