基于Pinecone+LangChain+Node.js的知识库聊天优化:解决上下文不足问题
问题分析
你的核心问题是虽然调用相似性搜索时指定了返回数量,但最终只使用了第一条搜索结果的内容作为上下文,导致ChatGPT获取的参考信息不足,回答质量达不到预期。
解决方案
1. 整合所有搜索结果到上下文
将所有搜索到的相关文档片段拼接成完整的上下文,而不是仅取第一条。这样ChatGPT能获取更全面的信息支撑回答。
2. 调整默认返回数量
把ranking参数的默认值从1改为更大的数值(比如3或5),确保默认情况下能获取足够的相关片段。
3. 增加空结果处理(可选)
添加对搜索结果为空的判断,避免因无上下文导致的错误回答。
修改后的代码片段
重点修改searchSimilarDocument函数中的上下文拼接和默认值部分:
export const searchSimilarDocument = async (req: Request, res: Response) => { const role: string = req.body.role as string; const query: string = req.body.query as string; // 将默认ranking从1改为3,可根据实际需求调整 const ranking: number = (req.body.ranking as number) || 3; const client = new PineconeClient(); await client.init({ apiKey: process.env.PINECONE_API_KEY ?? '', environment: process.env.PINECONE_ENVIRONMENT ?? '', }); const pineconeIndex = client.Index(process.env.PINECONE_INDEX ?? ''); const vectorStore = await PineconeStore.fromExistingIndex( new OpenAIEmbeddings(), { pineconeIndex: pineconeIndex, namespace: process.env.PINECONE_NAMESPACE ?? '', textKey: 'markdown' } ); try { const results = await vectorStore.similaritySearch(query, ranking, { role: role, }); // 处理无搜索结果的情况 if (results.length === 0) { return res.json({ content: "未找到相关信息,请尝试调整问题表述。" }); } // 拼接所有搜索结果的内容作为上下文 const context = results.map(item => item.pageContent).join('\n\n---------------------/\n\n'); const chat = new ChatOpenAI({ temperature: 0 }); const response = await chat.call([ new SystemChatMessage( `我们提供了以下上下文信息: ---------------------/ ${context} ---------------------/ 请基于这些信息,用我提问的语言回答我的问题。` ), new HumanChatMessage(query), ]); res.json(response); } catch (err) { console.error('Query Error: ' + err); res.status(500).json({ error: '查询过程中发生错误' }); } };
额外优化建议
- 调整文本分割策略:如果当前
MarkdownTextSplitter分割的片段过长或过短,可自定义chunkSize和chunkOverlap参数,让分割后的片段更适配语义搜索。 - 使用MMR搜索:LangChain的PineconeStore支持
maxMarginalRelevanceSearch方法,能在保证相关性的同时提升结果多样性,避免重复内容干扰。 - 添加相关性过滤:对搜索结果的相关性打分做阈值过滤,只保留分数达标的片段,减少无关信息对回答的影响。
内容的提问来源于stack exchange,提问作者melodyxpot
相关产品推荐
相关产品推荐

