You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Langchain RetrievalQA调用Pinecone时触发BaseRetriever验证错误求助

问题修复方案

错误根源

这个ValidationError是因为传入RetrievalQA的retriever对象没有正确实现BaseRetriever的所有抽象方法(_aget_relevant_documents和_get_relevant_documents),大概率是LangChain生态依赖版本不兼容,或者VectorStore初始化流程有误导致的。

具体修复步骤

1. 升级兼容依赖版本

先确保所有相关库是最新兼容版本,执行以下命令:

pip install --upgrade langchain langchain-pinecone pinecone-client

2. 修正VectorStore初始化逻辑

你的代码里重复初始化了docsearch(先调用from_texts又调用from_existing_index),保留其中一种即可,同时要确保embeddings是合法的Embedding实例(比如本地模型或第三方Embedding服务)。

3. 正确创建Retriever实例

提前将Retriever赋值为变量,避免在RetrievalQA.from_chain_type参数中直接调用方法,确保返回的是合法实例:

retriever = docsearch.as_retriever(search_kwargs={'k': 2})

完整修正后的代码示例

from langchain_pinecone import PineconeVectorStore
from pinecone import Pinecone
from langchain.chains import RetrievalQA
from langchain.prompts import PromptTemplate
from langchain.llms import CTransformers
# 补充embeddings的定义,比如用本地的SentenceTransformerEmbeddings
from langchain.embeddings import SentenceTransformerEmbeddings

# 初始化Pinecone
pc = Pinecone(api_key="1e094edb-8730-46a7-8178-615a08ca303b")
index = pc.Index("medical-chatbot")

# 初始化Embeddings(根据你的实际需求选择)
embeddings = SentenceTransformerEmbeddings(model_name="all-MiniLM-L6-v2")

# 二选一:如果是首次创建索引用from_texts,否则用from_existing_index
# 首次创建(注释掉下面的from_existing_index)
# docsearch = PineconeVectorStore.from_texts([t.page_content for t in text_chunks], embeddings, index_name='medical-chatbot')

# 加载已存在的索引
index_name = "medical-chatbot"
docsearch = PineconeVectorStore.from_existing_index(index_name, embeddings)

# 测试相似性搜索
query = "What are Salivary Gland Disease"
docs = docsearch.similarity_search(query, k=3)
print("Result", docs)

# 定义Prompt模板
prompt_template = """
Use the following pieces of information to answer the user's question.
If you don't know the answer, just say that you don't know, don't try to make up an answer.

Context: {context}
Question: {question}

Only return the helpful answer below and nothing else.
Helpful answer:
"""
PROMPT = PromptTemplate(template=prompt_template, input_variables=["context", "question"])
chain_type_kwargs = {"prompt": PROMPT}

# 加载LLM模型
llm = CTransformers(model="../model/llama-2-7b-chat.ggmlv3.q4_0.bin",
                    model_type="llama",
                    config={'max_new_tokens': 512,
                            'temperature': 0.8})

# 创建Retriever实例
retriever = docsearch.as_retriever(search_kwargs={'k': 2})

# 构建RetrievalQA链
qa = RetrievalQA.from_chain_type(
    llm=llm,
    chain_type="stuff",
    retriever=retriever,
    return_source_documents=True,
    chain_type_kwargs=chain_type_kwargs
)

额外注意点

  • 确保embeddings变量是正确初始化的Embedding实例,不能为None或未定义。
  • 如果使用本地Embedding模型,需要提前安装对应的依赖(比如sentence-transformers)。

内容的提问来源于stack exchange,提问作者Fakhruddin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 16:33:27