You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

LangChain自定义Prompt无效:Excel查询需返回JSON格式未生效

问题:LangChain自定义Prompt无法返回JSON格式结果

我正在使用LangChain读取Excel数据并进行查询响应,功能运行正常,但希望返回JSON格式的结果,为此添加了自定义Prompt,但该Prompt完全不起作用。以下是我的test_excel函数代码:

def test_excel(self, query="give me price for samsung"):
    loader = UnstructuredExcelLoader("assets/Master.xlsx", mode="elements")
    documents = loader.load()
    text_splitter = CharacterTextSplitter(chunk_size=1000, chunk_overlap=0)
    documents = text_splitter.split_documents(documents)

    embeddings = OpenAIEmbeddings()
    vectorstore = Chroma.from_documents(documents, embeddings)
    memory = ConversationBufferMemory(memory_key="chat_history", return_messages=True)

    chat_history = []

    custom_template = """ 
       Provide your answer in JSON format with the following keys: 
           Brand, Screen Size, Box size, price, Recommendation
    Chat History:
    {chat_history}
    Follow Up Input: {question}
    Standalone question:"""

    CUSTOM_QUESTION_PROMPT = PromptTemplate.from_template(custom_template)

    qa = ConversationalRetrievalChain.from_llm(OpenAI(temperature=0), vectorstore.as_retriever(),
                                               condense_question_prompt=CUSTOM_QUESTION_PROMPT,
                                               memory=memory
                                               )
    result = qa({"question": query, "chat_history": chat_history})
   
    return result

错误原因

你把自定义Prompt用在了错误的位置:

  • condense_question_prompt的作用是将对话历史和当前问题合并成一个独立的、无需上下文就能理解的问题,它只影响向向量库检索时用的查询语句,完全不控制最终回答的格式。你把要求返回JSON的模板放在这里,相当于让AI把问题转成JSON,这显然不是你要的效果。

修正方案

要控制最终回答的格式,需要修改的是文档内容与问题结合生成回答的Prompt。可以通过创建自定义的StuffDocumentsChain来实现,传入自定义的回答模板:

def test_excel(self, query="give me price for samsung"):
    loader = UnstructuredExcelLoader("assets/Master.xlsx", mode="elements")
    documents = loader.load()
    text_splitter = CharacterTextSplitter(chunk_size=1000, chunk_overlap=0)
    documents = text_splitter.split_documents(documents)

    embeddings = OpenAIEmbeddings()
    vectorstore = Chroma.from_documents(documents, embeddings)
    memory = ConversationBufferMemory(memory_key="chat_history", return_messages=True)

    chat_history = []

    # 自定义回答模板:明确要求返回指定格式的JSON
    answer_template = """
    根据提供的文档内容和对话历史,回答用户的问题,必须以JSON格式返回,包含以下必填键:Brand, Screen Size, Box size, price, Recommendation。
    如果某些信息无法从文档中获取,对应字段值设为"未知"。

    文档内容:
    {context}

    对话历史:
    {chat_history}

    用户问题:{question}

    你的JSON回答:
    """
    ANSWER_PROMPT = PromptTemplate.from_template(answer_template)

    # 创建自定义的文档合并链,绑定回答模板
    llm = OpenAI(temperature=0)
    combine_docs_chain = StuffDocumentsChain.from_llm(
        llm=llm,
        document_variable_name="context",
        prompt=ANSWER_PROMPT
    )

    # 初始化对话检索链,传入自定义的文档合并链
    qa = ConversationalRetrievalChain(
        retriever=vectorstore.as_retriever(),
        combine_docs_chain=combine_docs_chain,
        memory=memory
    )

    result = qa({"question": query, "chat_history": chat_history})
   
    return result

额外说明

  • 模板里明确了缺失信息的处理方式,避免AI因找不到对应数据而拒绝返回JSON格式内容。
  • 设置temperature=0可以让LLM输出更稳定,减少格式混乱的概率。

内容的提问来源于stack exchange,提问作者junaidp

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 23:51:19