You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Langchain的PGVector多集合查询:能否传入多个auth_id?

解决PGVector多用户集合查询的问题

直接在collection_name参数中传入多个auth_id是不可行的——PGVector的该参数仅支持单个字符串作为集合名称,无法同时指定多个集合。不过可以通过以下两种方式实现多用户文件的查询需求:

方案1:创建多个PGVector实例,合并查询结果

循环遍历所有auth_id,为每个auth_id创建独立的PGVector实例,分别执行查询后合并结果:

from langchain.vectorstores.pgvector import PGVector
from langchain.embeddings.openai import OpenAIEmbeddings

auth_ids = ["123", "456"]
openai_key = "your_api_key"
CONNECTION_STRING = "your_connection_string"
MAX_RETRIES = 3
DISTANCE_STRATEGY = "cosine"

def get_multi_collection_results(query, auth_ids):
    all_results = []
    embedding_func = OpenAIEmbeddings(openai_api_key=openai_key, max_retries=MAX_RETRIES)
    
    for auth_id in auth_ids:
        vector_store = PGVector(
            embedding_function=embedding_func,
            collection_name=auth_id,
            connection_string=CONNECTION_STRING,
            distance_strategy=DISTANCE_STRATEGY
        )
        # 可根据需求选择similarity_search/similarity_search_with_score等查询方法
        results = vector_store.similarity_search(query, k=5)
        all_results.extend(results)
    
    # 可选:对合并后的结果去重或重新排序
    return all_results

方案2:单集合+用户标识字段,通过过滤实现多用户查询

修改数据存储逻辑,将所有用户的向量数据存入同一个集合,同时为每条数据添加auth_id字段作为标识,查询时通过filter参数指定多个auth_id:

存储数据时添加auth_id元数据

from langchain.docstore.document import Document
from langchain.vectorstores.pgvector import PGVector

embedding_func = OpenAIEmbeddings(openai_api_key=openai_key, max_retries=MAX_RETRIES)

# 示例文档,每条携带对应auth_id
documents = [
    Document(page_content="用户1的文档内容", metadata={"auth_id": "123"}),
    Document(page_content="用户2的文档内容", metadata={"auth_id": "456"})
]

# 将所有文档存入统一集合(比如命名为"all_users")
vector_store = PGVector.from_documents(
    documents=documents,
    embedding=embedding_func,
    collection_name="all_users",
    connection_string=CONNECTION_STRING
)

查询时过滤多个auth_id

def get_filtered_results(query, auth_ids):
    vector_store = PGVector(
        embedding_function=embedding_func,
        collection_name="all_users",
        connection_string=CONNECTION_STRING,
        distance_strategy=DISTANCE_STRATEGY
    )
    # 使用$in操作符匹配多个auth_id
    filter_condition = {"auth_id": {"$in": auth_ids}}
    results = vector_store.similarity_search(query, k=5, filter=filter_condition)
    return results

两种方案对比

  • 方案1无需修改现有存储结构,适合快速实现,但多次查询会增加数据库交互次数,数据量大时效率较低。
  • 方案2仅需一次查询,效率更高,便于统一管理数据,但需要调整原有存储流程,确保存入数据时携带auth_id元数据。

内容的提问来源于stack exchange,提问作者Kajol Mehta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 23:03:28