You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于LangChain/pgvector的语义相似文档检索优化及RAG应用问询

文档检索优化与RAG缩写学习方案咨询

背景与需求

我有一批包含标签(tags)、位置(location)、**文本(text)**三个属性的文档,目前用LangChain+pgvector+Embeddings对所有属性做索引,效果还不错。但想找更优方案:

  • 需要检索带有特定标签和位置、文本表述差异极大但语义相同的文档,打算用Embeddings+向量数据库实现;
  • 另外,能不能借助RAG(检索增强生成)让大语言模型(LLM)学习它原本不知道的常用缩写?

当前实现代码(已翻译注释)

import pandas as pd

from langchain_core.documents import Document
from langchain_postgres import PGVector
from langchain_postgres.vectorstores import PGVector
from langchain_openai.embeddings import OpenAIEmbeddings

# 数据库连接字符串
connection = "postgresql+psycopg://langchain:langchain@localhost:5432/langchain"
# 初始化OpenAI嵌入模型
embeddings = OpenAIEmbeddings(model="text-embedding-3-small")
# 向量集合名称
collection_name = "notas_v0"

# 初始化PGVector向量存储
vectorstore = PGVector(
    embeddings=embeddings,
    collection_name=collection_name,
    connection=connection,
    use_jsonb=True,
)


# --- 索引构建部分(已注释)---

# # 读取CSV数据
# df = pd.read_csv("notes.csv")
# # 去除空值
# df = df.dropna()  # 可添加.head(10000)限制处理数量
# # 处理标签列,拆分字符串为标签列表
# df["tags"] = df["tags"].apply(
#     lambda x: [tag.strip() for tag in x.split(",") if tag.strip()]
# )

# # 提取各字段数据
# long_texts = df["Texto Longo"].tolist()  # 长文本列
# wc = df["Centro Trabalho Responsável"].tolist()  # 负责工作中心列
# notes = df["Nota"].tolist()  # 备注列
# tags = df["tags"].tolist()

# # 构建LangChain Document对象列表
# documents = list(
#     map(
#         lambda x: Document(
#             page_content=x[0], metadata={"wc": x[1], "note": x[2], "tags": x[3]}
#         ),
#         zip(long_texts, wc, notes, tags),
#     )
# )

# # 批量添加文档到向量库(每100条一批)
# print(
#     [
#         vectorstore.add_documents(documents=documents[i : i + 100])
#         for i in range(0, len(documents), 100)
#     ]
# )
# print("索引构建完成。")

# --- 索引构建结束 ---

# --- 查询部分 ---

# 执行相似度检索,带过滤条件和相关性分数
result = vectorstore.similarity_search_with_relevance_scores(
    "EVTD202301222707",  # 查询文本
    filter={"note": {"$in": ["15310116"]}, "tags": {"$in": ["abcd", "xyz"]}},  # 过滤条件:特定备注和标签
    k=10, # 返回结果数量上限
)

# --- 查询结束 ---

内容的提问来源于stack exchange,提问作者Rodrigo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 18:27:26