基于PDF训练的LLM答非所问,如何实现上下文外问题拒答?
解决LLM回答超出PDF上下文范围问题的方案
1. 优化提示模板,明确回答规则
LLM的无中生有回答大多源于缺乏明确指令约束,需在提示中强制模型仅基于给定上下文作答,超出范围则返回指定内容。
针对lmsys/fastchat-t5-3b-v1.0的提示模板
question_t5_template = """ context: {context} question: {question} Instructions: Answer the question ONLY using the information provided in the context. If the question cannot be answered using the context, respond with one of "您的问题超出训练数据范围", "I don't know", "out of context". answer: """ QUESTION_T5_PROMPT = PromptTemplate( template=question_t5_template, input_variables=["context", "question"] ) qa.combine_documents_chain.llm_chain.prompt = QUESTION_T5_PROMPT
针对falcon-7b-instruct的提示模板
Falcon更适配对话式指令格式,调整模板如下:
question_falcon_template = """ ### Instruction: Answer the question ONLY using the information provided in the following context. If the question cannot be answered using the context, respond with one of "您的问题超出训练数据范围", "I don't know", "out of context". ### Context: {context} ### Question: {question} ### Answer: """ QUESTION_FALCON_PROMPT = PromptTemplate( template=question_falcon_template, input_variables=["context", "question"] ) qa.combine_documents_chain.llm_chain.prompt = QUESTION_FALCON_PROMPT
2. 添加检索结果校验,提前拦截无关问题
如果检索到的文档与问题相似度极低,说明问题不在PDF上下文范围内,可直接返回指定内容,无需调用LLM。
修改answer_query函数:
def answer_query(self, question: str) -> str: """Answer the question""" answer_dict = self.qa({"query": question}) source_docs = answer_dict.get("source_documents", []) # 检查检索结果的相似度得分(需确保RetrievalQA配置返回得分) # 不同检索实现的得分字段可能不同,比如`score`或`similarity_score` threshold = 0.5 # 根据实际效果调整阈值 has_relevant_docs = any(doc.metadata.get("score", 0) >= threshold for doc in source_docs) if not has_relevant_docs or not source_docs: return "您的问题超出训练数据范围" answer = answer_dict["result"] # 可选:额外校验LLM返回内容是否与上下文相关(比如用嵌入模型计算相似度) return answer
3. 额外优化建议
- 微调检索阈值:根据PDF内容和问题类型调整相似度阈值,平衡召回率与准确率。
- 强化指令遵循:对T5设置
temperature=0、对Falcon设置do_sample=False,减少模型生成随机性,提升指令执行度。
内容的提问来源于stack exchange,提问作者Mukilan
相关产品推荐
相关产品推荐

