You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Pinecone+LangChain+Node.js的知识库聊天优化:解决上下文不足问题

问题分析

你的核心问题是虽然调用相似性搜索时指定了返回数量,但最终只使用了第一条搜索结果的内容作为上下文,导致ChatGPT获取的参考信息不足,回答质量达不到预期。

解决方案

1. 整合所有搜索结果到上下文

将所有搜索到的相关文档片段拼接成完整的上下文,而不是仅取第一条。这样ChatGPT能获取更全面的信息支撑回答。

2. 调整默认返回数量

把ranking参数的默认值从1改为更大的数值(比如3或5),确保默认情况下能获取足够的相关片段。

3. 增加空结果处理(可选)

添加对搜索结果为空的判断,避免因无上下文导致的错误回答。

修改后的代码片段

重点修改searchSimilarDocument函数中的上下文拼接和默认值部分:

export const searchSimilarDocument = async (req: Request, res: Response) => {

  const role: string = req.body.role as string;
  const query: string = req.body.query as string;
  // 将默认ranking从1改为3,可根据实际需求调整
  const ranking: number = (req.body.ranking as number) || 3;

  const client = new PineconeClient();
  await client.init({
    apiKey: process.env.PINECONE_API_KEY ?? '',
    environment: process.env.PINECONE_ENVIRONMENT ?? '',
  });

  const pineconeIndex = client.Index(process.env.PINECONE_INDEX ?? '');
  const vectorStore = await PineconeStore.fromExistingIndex(
    new OpenAIEmbeddings(),
    { 
      pineconeIndex: pineconeIndex,
      namespace: process.env.PINECONE_NAMESPACE ?? '',
      textKey: 'markdown'
    }
  );

  try {
    const results = await vectorStore.similaritySearch(query, ranking, {
      role: role,
    });

    // 处理无搜索结果的情况
    if (results.length === 0) {
      return res.json({ content: "未找到相关信息,请尝试调整问题表述。" });
    }

    // 拼接所有搜索结果的内容作为上下文
    const context = results.map(item => item.pageContent).join('\n\n---------------------/\n\n');

    const chat = new ChatOpenAI({ temperature: 0 });

    const response = await chat.call([
      new SystemChatMessage(
        `我们提供了以下上下文信息:

        ---------------------/

        ${context}
        ---------------------/

        请基于这些信息,用我提问的语言回答我的问题。`
      ),
      new HumanChatMessage(query),
    ]);

    res.json(response);
  } catch (err) {
    console.error('Query Error: ' + err);
    res.status(500).json({ error: '查询过程中发生错误' });
  }
};
额外优化建议
  • 调整文本分割策略:如果当前MarkdownTextSplitter分割的片段过长或过短,可自定义chunkSize和chunkOverlap参数,让分割后的片段更适配语义搜索。
  • 使用MMR搜索:LangChain的PineconeStore支持maxMarginalRelevanceSearch方法,能在保证相关性的同时提升结果多样性,避免重复内容干扰。
  • 添加相关性过滤:对搜索结果的相关性打分做阈值过滤,只保留分数达标的片段,减少无关信息对回答的影响。

内容的提问来源于stack exchange,提问作者melodyxpot

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 13:12:55