You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Qdrant向量库元数据与上下文一同发送给OpenAI大模型?

解决LangChain RetrievalQA返回元数据(文件名)的问题

核心思路

要让LLM在回答中包含简历的来源文件名,需要完成两个关键操作:

  • 将文档的元数据(source字段)与正文内容拼接成完整上下文,传递给LLM
  • 在提示模板中明确要求LLM在回答中引用对应的来源文件名

具体实现代码

修改你的run_llm函数,通过自定义文档格式化模板和问答提示模板,将元数据注入上下文并引导LLM返回来源信息:

from langchain.prompts import PromptTemplate
from langchain.chat_models import ChatOpenAI
from langchain.chains import RetrievalQA

def run_llm(query: str):
    # 自定义文档格式化模板:将每个文档的内容和source元数据组合
    document_prompt = PromptTemplate(
        input_variables=["page_content", "source"],
        template="简历内容:{page_content}\n来源文件:{source}\n---"
    )
    
    # 自定义问答提示模板:明确要求LLM返回来源文件名
    qa_prompt = PromptTemplate(
        template="请根据以下提供的简历上下文回答问题,回答时必须附上对应的来源文件名。\n\n上下文:\n{context}\n\n问题:{question}\n\n回答:",
        input_variables=["context", "question"]
    )
    
    chat = ChatOpenAI(verbose=True, temperature=0)
    qa = RetrievalQA.from_chain_type(
        llm=chat,
        chain_type="stuff",
        retriever=qdrant.as_retriever(),
        return_source_documents=True,
        # 将自定义模板传入chain配置
        chain_type_kwargs={
            "prompt": qa_prompt,
            "document_prompt": document_prompt
        }
    )
    return qa({"query": query})

print(run_llm(query="Which unique resume has the longest working experience as accountant?"))

关键说明

  1. document_prompt:负责将每个检索到的文档的page_content和source元数据格式化为LLM可识别的上下文条目,确保元数据被包含在输入中。
  2. qa_prompt:明确告知LLM需要在回答中附上来源文件名,引导模型输出符合要求的结果。
  3. 保持return_source_documents=True可以同时获取原始来源文档,方便后续验证。

内容的提问来源于stack exchange,提问作者Ruby Rain

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 02:57:29