Langchain中RetrievalQA.from_chain_type报错ValidationError求助
解决LangChain中RetrievalQA的ValidationError问题
问题场景
运行以下LangChain代码构建RetrievalQA链时,触发了ValidationError:
from langchain_google_genai import GoogleGenerativeAIEmbeddings google_generative_ai_Embeddings = GoogleGenerativeAIEmbeddings(model="models/embedding-001", google_api_key=api_key) from langchain_community.vectorstores import Chroma vectordb = Chroma.from_documents(data, embedding=google_generative_ai_Embeddings, persist_directory='./chromadb') retriever_google = vectordb.as_retriever(score_threshold = 0.7) from langchain.prompts import PromptTemplate prompt_template = """Given the following context and a question, generate an answer based on this context only. In the answer try to provide as much text as possible from "response" section in the source document context without making much changes. If the answer is not found in the context, kindly state "I don't know." Don't try to make up an answer. CONTEXT: {context} QUESTION: {question}""" PROMPT = PromptTemplate( template=prompt_template, input_variables=["context", "question"] ) chain_type_kwargs = {"prompt": PROMPT} from langchain.chains import RetrievalQA chain_type = "stuff" chain = RetrievalQA.from_chain_type(llm=llm, chain_type=chain_type, retriever=retriever_google, input_key="query", return_source_documents=True, chain_type_kwargs=chain_type_kwargs)
错误信息如下:
ValidationError Traceback (most recent call last) ~\AppData\Local\Temp\ipykernel_4328\826232400.py in <module> 20 chain_type = "stuff" 21 ---> 22 chain = RetrievalQA.from_chain_type(llm=llm, 23 chain_type=chain_type, 24 retriever=retriever_google, ~\AppData\Roaming\Python\Python39\site-packages\langchain\chains\retrieval_qa\base.py in from_chain_type(cls, llm, chain_type, chain_type_kwargs, **kwargs) 98 """Load chain from chain type.""" 99 _chain_type_kwargs = chain_type_kwargs or {} --> 100 combine_documents_chain = load_qa_chain( 101 llm, chain_type=chain_type, **_chain_type_kwargs 102 ) ~\AppData\Roaming\Python\Python39\site-packages\langchain\chains\question_answering\__init__.py in load_qa_chain(llm, chain_type, verbose, callback_manager, **kwargs) 247 f"Should be one of {loader_mapping.keys()}" 248 ) --> 249 return loader_mapping[chain_type]( 250 llm, verbose=verbose, callback_manager=callback_manager, **kwargs 251 ) ~\AppData\Roaming\Python\Python39\site-packages\langchain\chains\question_answering\__init__.py in _load_stuff_chain(llm, prompt, document_variable_name, verbose, callback_manager, callbacks, **kwargs) 71 ) -> StuffDocumentsChain: 72 _prompt = prompt or stuff_prompt.PROMPT_SELECTOR.get_prompt(llm) --> 73 llm_chain = LLMChain( 74 llm=llm, 75 prompt=_prompt, ~\AppData\Roaming\Python\Python39\site-packages\langchain\load\serializable.py in __init__(self, **kwargs) 73 74 def __init__(self, **kwargs: Any) -> None: --> 75 super().__init__(**kwargs) 76 self._lc_kwargs = kwargs 77 ~\AppData\Roaming\Python\Python39\site-packages\pydantic\v1\main.py in __init__(__pydantic_self__, **data) 339 values, fields_set, validation_error = validate_model(__pydantic_self__.__class__, data) 340 if validation_error: --> 341 raise validation_error 342 try: 343 object_setattr(__pydantic_self__, '__dict__', values) ValidationError: 1 validation error for LLMChain llm Can't instantiate abstract class BaseLanguageModel with abstract methods agenerate_prompt, apredict, apredict_messages, generate_prompt, invoke, predict, predict_messages (type=type_error)
错误原因
代码中的llm变量未被正确实例化:
- 仅引用了
llm但未定义,或错误使用LangChain的抽象基类BaseLanguageModel而非具体模型实现 - RetrievalQA需要接收已实例化的大语言模型对象,而非抽象类或未定义变量
修复步骤
- 导入对应LLM的实现类(这里搭配你使用的Google Embedding模型,选择Google Gemini)
- 用你的Google API Key实例化LLM对象
- 将实例化后的
llm传入RetrievalQA
修复后的完整代码
from langchain_google_genai import GoogleGenerativeAIEmbeddings, GoogleGenerativeAI # 实例化Embedding模型 google_generative_ai_Embeddings = GoogleGenerativeAIEmbeddings(model="models/embedding-001", google_api_key=api_key) # 实例化LLM(关键补充步骤) llm = GoogleGenerativeAI(model="gemini-pro", google_api_key=api_key, temperature=0.0) from langchain_community.vectorstores import Chroma vectordb = Chroma.from_documents(data, embedding=google_generative_ai_Embeddings, persist_directory='./chromadb') retriever_google = vectordb.as_retriever(score_threshold = 0.7) from langchain.prompts import PromptTemplate prompt_template = """Given the following context and a question, generate an answer based on this context only. In the answer try to provide as much text as possible from "response" section in the source document context without making much changes. If the answer is not found in the context, kindly state "I don't know." Don't try to make up an answer. CONTEXT: {context} QUESTION: {question}""" PROMPT = PromptTemplate( template=prompt_template, input_variables=["context", "question"] ) chain_type_kwargs = {"prompt": PROMPT} from langchain.chains import RetrievalQA chain_type = "stuff" chain = RetrievalQA.from_chain_type(llm=llm, chain_type=chain_type, retriever=retriever_google, input_key="query", return_source_documents=True, chain_type_kwargs=chain_type_kwargs)
注意事项
- 确保
api_key变量已正确设置为你的Google API密钥 - 根据需求调整LLM参数,比如
temperature控制生成内容的随机性 - 如果使用其他LLM(如OpenAI、Anthropic等),替换对应的导入和实例化代码即可
内容的提问来源于stack exchange,提问作者Monidipa Das
相关产品推荐
相关产品推荐

