Azure部署的RetrievalQA无法返回源文档,本地运行正常求助
Azure部署RetrievalQA.from_chain_type时source_documents为空,本地运行正常的排查方案
问题详情
本地运行RetrievalQA时能正常返回源文档列表,但部署到Azure后,所有查询的source_documents均为空,且返回结果内容与本地不一致:
Azure返回结果
{'query': 'Please provide rationale.. ', 'result': 'Creative Output: The rationale behind the process is to ensure that all actions taken are based on sound reasoning and logic, and are aligned with the overall objectives and goals of the organization. This helps to minimize risks, optimize resources, and improve performance.', 'source_documents': []}
本地返回结果
{'query': 'Please provide rationale.. ', 'result': 'Extracted Output: Could you please provide more context or specify which process or policy you are referring to? Without additional information, it is difficult to provide a concise response.', 'source_documents': [Document(page_content='Please provide evidence that this occurred in the last twelve months.', metadata={'source_page': 'Third Party-Vendor Audit Process Version 378'}), Document(page_content='Please provide evidence that this occurred in the last twelve months.', metadata={'source_page': 'Third Party-Vendor Audit Process Version 371'}), Document(page_content='Please Provide evidence such as emails or other notifications showing periodic access reviews.', metadata={'source_page': 'Third Party-Vendor Audit Process Version 352'}), Document(page_content='Please provide evidence that this has occurred within the last twelve months.', metadata={'source_page': 'Third Party-Vendor Audit Process Version 331'})]}
环境差异:本地向量数据库存储在本地目录,Azure环境下存储在GitHub的vdb_langchain_small目录,使用的代码完全一致:
PROMPT=PromptTemplate(template=Prompt_Template,input_variables=["context","question"]) chain_type_kwargs={"prompt":PROMPT} qa_chain = RetrievalQA.from_chain_type(llm=self.llm_open, chain_type="stuff", retriever=self.retriever, return_source_documents=True, chain_type_kwargs=chain_type_kwargs, verbose=True) qa_results = {}
排查方向与解决方案
1. 验证向量库加载路径与访问权限
- Azure环境中,直接读取GitHub远程目录的向量库文件可能存在权限或路径解析问题。需确认向量库文件是否已同步到Azure本地环境,而非直接引用远程GitHub路径。
- 打印retriever加载的向量库路径,确认路径有效性:
# 根据使用的向量库类型调整,例如FAISS print(self.retriever.vectorstore._path) - 若使用GitHub仓库存储向量库,需确保Azure环境已克隆该仓库,且代码中使用的是克隆后的本地绝对路径。
2. 检查Embedding模型一致性
- 本地与Azure环境使用的Embedding模型必须完全一致(包括模型名称、部署版本、参数配置)。如果Embedding模型不同,向量相似度匹配结果会完全失效,导致retriever无法检索到文档。
- 例如,本地使用OpenAI的
text-embedding-ada-002,Azure上需部署相同模型并使用对应API调用。
3. 统一依赖版本
- 导出本地环境的依赖清单(
requirements.txt),在Azure环境中安装完全相同版本的LangChain及相关依赖(如FAISS、Chroma等)。版本差异可能导致API行为不一致,进而影响retriever的文档返回。 - 示例依赖清单:
langchain==0.1.10 faiss-cpu==1.7.4 openai==1.13.3
4. 单独调试Retriever功能
- 脱离RetrievalQA链,单独测试retriever的检索能力:
docs = self.retriever.get_relevant_documents("Please provide rationale..") print(docs)- 若返回为空:说明向量库未正确加载或Embedding匹配失败,重点排查向量库路径、文件完整性及Embedding模型配置。
- 若返回正常文档:说明问题出在RetrievalQA链的配置环节,需检查
chain_type_kwargs中的prompt模板是否正确传递上下文,或链的参数是否存在环境差异。
内容的提问来源于stack exchange,提问作者Mohit
相关产品推荐
相关产品推荐

