如何让LangChain聊天机器人仅基于自定义知识库作答?
解决LangChain聊天机器人仅依托自定义知识库作答的问题
核心问题分析
当前聊天机器人会调用GPT-3.5自身知识库回答未涵盖在自定义docs中的问题,根源在于:
- RetrievalQA未严格约束LLM只能使用检索到的文档内容
- 未对检索结果的相关性做判断,导致无关文档被传入LLM生成内容
- Agent系统提示的约束性不足,允许LLM直接使用外部知识
具体修改方案
1. 自定义严格约束的PromptTemplate
给RetrievalQA设置专属提示词,强制LLM仅基于检索到的文档内容回答,禁止使用外部知识,同时明确无相关内容时的固定回复话术。
qa_prompt = PromptTemplate( template="""Use the following pieces of context to answer the user's question. If you don't find the answer in the context, just say "Sorry, I wasn't provided with information that answers your question" and nothing else. Do NOT use any knowledge outside of the provided context. Context: {context} Question: {question} Answer:""", input_variables=["context", "question"] )
2. 调整RetrievalQA配置,启用相关性过滤
设置检索器的search_kwargs限制返回最相关的文档数量,同时将自定义提示词传入RetrievalQA,确保LLM严格遵循规则。
# 调整检索器,返回最相关的3条文档(可根据需求调整k值) retriever = vectorstore.as_retriever(search_kwargs={"k": 3}) # 初始化RetrievalQA时绑定自定义提示词 qa = RetrievalQA.from_chain_type( llm=llm, chain_type="stuff", retriever=retriever, verbose=True, return_source_documents=True, chain_type_kwargs={"prompt": qa_prompt} # 传入自定义提示词 )
3. 自定义工具处理逻辑,过滤无效检索结果
新增包装函数,检查RetrievalQA返回的结果是否为预设的无答案回复,若为真则清空源文档列表,避免返回无意义内容。
def custom_qa_tool(inputs): result = qa({"query": inputs}) # 判断是否为无答案回复,清空无效源文档 if result["result"] == "Sorry, I wasn't provided with information that answers your question": result["source_documents"] = [] return result
同步修改工具定义,强化工具的必要性:
tools = [ Tool( name="doc_search_tool", func=custom_qa_tool, description=( "This tool is used to retrieve information from the knowledge base. " "Use this tool for ALL questions, do NOT answer without using this tool." ) ) ]
4. 强化Agent系统提示的约束性
更新系统提示,明确要求Agent必须先使用工具获取信息,禁止直接回答问题:
system_message = """ You are the XYZ bot. This is a conversation with a human. You MUST use the doc_search_tool to get information before answering any question. Answer ONLY based on the information provided by the doc_search_tool. If the tool returns that no information is available, just say "Sorry, I wasn't provided with information that answers your question" and nothing else. Do NOT make up answers or use any knowledge outside of the tool's output. """
完整修改后的代码
import os from langchain.embeddings.openai import OpenAIEmbeddings from langchain.vectorstores import FAISS from langchain.chains import RetrievalQA from langchain.memory import ConversationBufferMemory from langchain.chat_models import ChatOpenAI from langchain.prompts import PromptTemplate from langchain.agents import AgentExecutor, Tool, initialize_agent from langchain.agents.types import AgentType os.environ['OPENAI_API_KEY'] = '' # 强化约束的系统提示 system_message = """ You are the XYZ bot. This is a conversation with a human. You MUST use the doc_search_tool to get information before answering any question. Answer ONLY based on the information provided by the doc_search_tool. If the tool returns that no information is available, just say "Sorry, I wasn't provided with information that answers your question" and nothing else. Do NOT make up answers or use any knowledge outside of the tool's output. """ llm = ChatOpenAI( model_name="gpt-3.5-turbo", temperature=0 ) embeddings = OpenAIEmbeddings() docs = [ "Buildings are made out of brick", "Buildings are made out of wood", "Buildings are made out of stone", "Buildings are made out of atoms", "Buildings are made out of building materials", "Cars are made out of metal", "Cars are made out of plastic", ] vectorstore = FAISS.from_texts(docs, embeddings) # 调整检索器,返回最相关的3条文档 retriever = vectorstore.as_retriever(search_kwargs={"k": 3}) # 自定义严格约束的QA提示词 qa_prompt = PromptTemplate( template="""Use the following pieces of context to answer the user's question. If you don't find the answer in the context, just say "Sorry, I wasn't provided with information that answers your question" and nothing else. Do NOT use any knowledge outside of the provided context. Context: {context} Question: {question} Answer:""", input_variables=["context", "question"] ) qa = RetrievalQA.from_chain_type( llm=llm, chain_type="stuff", retriever=retriever, verbose=True, return_source_documents=True, chain_type_kwargs={"prompt": qa_prompt} ) # 自定义工具函数,处理无效检索结果 def custom_qa_tool(inputs): result = qa({"query": inputs}) if result["result"] == "Sorry, I wasn't provided with information that answers your question": result["source_documents"] = [] return result tools = [ Tool( name="doc_search_tool", func=custom_qa_tool, description=( "This tool is used to retrieve information from the knowledge base. " "Use this tool for ALL questions, do NOT answer without using this tool." ) ) ] memory = ConversationBufferMemory(memory_key="chat_history", input_key='input', return_messages=True, output_key='output') agent = initialize_agent( agent=AgentType.CHAT_CONVERSATIONAL_REACT_DESCRIPTION, tools=tools, llm=llm, memory=memory, return_source_documents=True, return_intermediate_steps=True, agent_kwargs={"system_message": system_message} )
修改效果说明
- 当提问内容在
docs中时,机器人会基于检索到的文档回答,并返回对应的源文档 - 当提问内容不在
docs中时,机器人会准确返回预设的无答案回复,且不会返回无意义的源文档 - 严格限制LLM只能使用自定义知识库内容,不会调用GPT-3.5自身的外部知识
内容的提问来源于stack exchange,提问作者Blue Cheese
相关产品推荐
相关产品推荐

