如何将Qdrant向量库元数据与上下文一同发送给OpenAI大模型?
解决LangChain RetrievalQA返回元数据(文件名)的问题
核心思路
要让LLM在回答中包含简历的来源文件名,需要完成两个关键操作:
- 将文档的元数据(
source字段)与正文内容拼接成完整上下文,传递给LLM - 在提示模板中明确要求LLM在回答中引用对应的来源文件名
具体实现代码
修改你的run_llm函数,通过自定义文档格式化模板和问答提示模板,将元数据注入上下文并引导LLM返回来源信息:
from langchain.prompts import PromptTemplate from langchain.chat_models import ChatOpenAI from langchain.chains import RetrievalQA def run_llm(query: str): # 自定义文档格式化模板:将每个文档的内容和source元数据组合 document_prompt = PromptTemplate( input_variables=["page_content", "source"], template="简历内容:{page_content}\n来源文件:{source}\n---" ) # 自定义问答提示模板:明确要求LLM返回来源文件名 qa_prompt = PromptTemplate( template="请根据以下提供的简历上下文回答问题,回答时必须附上对应的来源文件名。\n\n上下文:\n{context}\n\n问题:{question}\n\n回答:", input_variables=["context", "question"] ) chat = ChatOpenAI(verbose=True, temperature=0) qa = RetrievalQA.from_chain_type( llm=chat, chain_type="stuff", retriever=qdrant.as_retriever(), return_source_documents=True, # 将自定义模板传入chain配置 chain_type_kwargs={ "prompt": qa_prompt, "document_prompt": document_prompt } ) return qa({"query": query}) print(run_llm(query="Which unique resume has the longest working experience as accountant?"))
关键说明
- document_prompt:负责将每个检索到的文档的
page_content和source元数据格式化为LLM可识别的上下文条目,确保元数据被包含在输入中。 - qa_prompt:明确告知LLM需要在回答中附上来源文件名,引导模型输出符合要求的结果。
- 保持
return_source_documents=True可以同时获取原始来源文档,方便后续验证。
内容的提问来源于stack exchange,提问作者Ruby Rain
相关产品推荐
相关产品推荐

