Langchain RetrievalQAChain返回0文档仍生成正确答案的技术疑问
问题描述
- 我原本认为RetrievalQAChain仅基于数据库返回的文档作答,且Prompt明确要求:若不知道答案,只需回复“不知道”,不得编造内容。但实际测试中,向量数据库返回0条文档时,模型却给出了正确答案,推测ChatGPT调用了自身知识库,想确认是否存在遗漏配置。
- 此外还有疑问:当PromptTemplate含相关文档时,如何验证模型是基于给定上下文还是自身知识生成回复?
代码实现
export async function POST(req: NextRequest) { try { const prompt = new PromptTemplate({ template: "Use the following pieces of context to answer the question at the end. If you don't know the answer, just say that you don't know, don't try to make up an answer.\n\nContext: {context}\n\nQuestion: {question}\nHelpful Answer:", inputVariables: ["context", "question"], }); {/*Rest of the code */} const vectorStoreRetriever = vectorStore.asRetriever(5); console.log( "vectorStoreRetriever", await vectorStoreRetriever._getRelevantDocuments(message) ); const model = new OpenAI({ streaming: true, temperature: 0, timeout: 60000, }); // Create a chain that uses the OpenAI LLM and HNSWLib vector store. const chain = new RetrievalQAChain({ combineDocumentsChain: loadQAStuffChain(model, { prompt }), retriever: vectorStore.asRetriever(), }); const results = await chain.call({ query: "What is NextJs", }); // return NextResponse.json({ splitDocuments }); return NextResponse.json({ results }); } catch (e: any) { console.log(e); return NextResponse.json({ error: e.message }, { status: 500 }); } }
测试结果

解决方案
一、解决无上下文时模型调用自身知识库的问题
当向量库返回0条文档时,{context}会被替换为空字符串,但OpenAI模型仍会依赖自身知识作答。要强制模型遵守Prompt要求,可通过以下两种方式处理:
- 优化Prompt模板,明确上下文为空时的回复规则:
const prompt = new PromptTemplate({ template: "Use the following pieces of context to answer the question at the end. If there is no context provided, or the context doesn't contain the answer, just say that you don't know, don't try to make up an answer.\n\nContext: {context}\n\nQuestion: {question}\nHelpful Answer:", inputVariables: ["context", "question"], });
- 提前检测文档数量,如果检索结果为空,直接返回“不知道”,避免调用LLM:
// 先获取相关文档 const relevantDocs = await vectorStoreRetriever._getRelevantDocuments("What is NextJs"); if (relevantDocs.length === 0) { return NextResponse.json({ results: { text: "不知道" } }); } // 再执行chain调用 const results = await chain.call({ query: "What is NextJs" });
二、验证模型是否基于给定上下文生成回复
有两种实用方法可以验证:
- 方法一:传入错误/虚构上下文
比如问“What is NextJs”,给上下文设置为“NextJs是一款Java后端框架”。如果模型输出和错误上下文一致,说明它基于给定内容作答;如果输出正确答案,说明调用了自身知识库。 - 方法二:添加唯一虚构标识
在上下文中加入只有你知道的虚构信息,比如问“What is NextJs的官方代号”,上下文写“NextJs的官方代号是北极星”。如果模型输出“北极星”,说明它使用了给定上下文;如果输出真实官方代号(或表示不知道),说明没用到给定内容。
内容的提问来源于stack exchange,提问作者Anni
相关产品推荐
相关产品推荐

