You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Node.js下LangChain+Pinecone RAG应用如何返回查询关联分块

解决Node.js LangChain RAG应用获取相似分块的问题

在Python的LangChain中,我们可以给检索器设置return_source_documents=True来获取查询所用的最相似分块,但Node.js环境里没有这个参数。要实现相同功能,只需调整Runnable序列的结构,同时保留检索到的原始文档对象即可。

修改思路

  1. 不直接将检索器结果用formatDocumentsAsString处理,先获取完整的Document数组
  2. 拆分处理逻辑:用检索到的文档生成上下文字符串,同时保留原始文档
  3. 让最终的Chain返回包含回答和来源文档的对象

修改后的完整代码

import { ChatOpenAI, OpenAIEmbeddings } from "@langchain/openai";
import { formatDocumentsAsString } from "langchain/util/document";
import { PromptTemplate } from "@langchain/core/prompts";
import {
  RunnableSequence,
  RunnablePassthrough,
} from "@langchain/core/runnables";
import { StringOutputParser } from "@langchain/core/output_parsers";

import { PineconeStore } from "@langchain/pinecone";
import { Pinecone } from '@pinecone-database/pinecone';

export default async function langchain() {
  const model = new ChatOpenAI({
    apiKey: process.env.NEXT_PUBLIC_OPENAI_API_KEY,
    modelName: "gpt-3.5-turbo",
    temperature: 0,
    streaming: true,
  });

  async function getVectorStore() {
    try {
      const client = new Pinecone({
        apiKey: process.env.NEXT_PUBLIC_PINECONE_API_KEY || 'mal'
      });

      const embeddings = new OpenAIEmbeddings({
        apiKey: process.env.NEXT_PUBLIC_OPENAI_API_KEY,
        batchSize: 1536,
        model: "text-embedding-ada-002",
      });

      const index = client.Index(process.env.NEXT_PUBLIC_PINECONE_INDEX_NAME || 'mal');

      const vectorStore = await PineconeStore.fromExistingIndex(embeddings, {
        pineconeIndex: index,
        namespace: process.env.NEXT_PUBLIC_PINECONE_ENVIRONMENT,
        textKey: 'text',
      });

      return vectorStore;
    } catch (error) {
      console.log('error ', error);
      throw new Error('Something went wrong while getting vector store!');
    }
  }

  const formatChatHistory = (chatHistory: [string, string][]) => {
    const formattedDialogueTurns = chatHistory.map(
      (dialogueTurn) => `Human: ${dialogueTurn[0]}\nAssistant: ${dialogueTurn[1]}`
    );
    return formattedDialogueTurns.join("\n");
  };

  const condenseQuestionTemplate = `Given the following conversation and a follow-up question, rephrase the follow-up question to be a standalone question, in its original language.

  Chat History:
  {chat_history}
  Follow Up Input: {question}
  Standalone question:`;
  const CONDENSE_QUESTION_PROMPT = PromptTemplate.fromTemplate(
    condenseQuestionTemplate
  );

  const answerTemplate = `Answer the question based only on the following context:
  {context}

  Question: {question}
  `;
  const ANSWER_PROMPT = PromptTemplate.fromTemplate(answerTemplate);

  type ConversationalRetrievalQAChainInput = {
    question: string;
    chat_history: [string, string][];
  };

  const standaloneQuestionChain = RunnableSequence.from([
    {
      question: (input: ConversationalRetrievalQAChainInput) => input.question,
      chat_history: (input: ConversationalRetrievalQAChainInput) =>
        formatChatHistory(input.chat_history),
    },
    CONDENSE_QUESTION_PROMPT,
    model,
    new StringOutputParser(),
  ]);

  const vectorStore = await getVectorStore();
  const retriever = vectorStore.asRetriever();

  // 修改核心部分:同时获取来源文档和生成上下文
  const answerChain = RunnableSequence.from([
    RunnablePassthrough.assign({
      // 先检索到原始文档
      sourceDocuments: retriever,
    }),
    RunnablePassthrough.assign({
      // 用原始文档生成上下文字符串
      context: (input) => formatDocumentsAsString(input.sourceDocuments),
    }),
    {
      context: (input) => input.context,
      question: (input) => input,
    },
    ANSWER_PROMPT,
    model,
    new StringOutputParser(),
    // 将回答和来源文档合并返回
    (answer, input) => ({
      answer,
      sourceDocuments: input.sourceDocuments,
    }),
  ]);

  const conversationalRetrievalQAChain =
    standaloneQuestionChain.pipe(answerChain);

  const result = await conversationalRetrievalQAChain.invoke({
    question: "Que es la SS?",
    chat_history: [],
  });

  console.log('Answer: ', result.answer);
  console.log('Source Documents: ', result.sourceDocuments);

  return result;
}

关键改动说明

  • 使用RunnablePassthrough.assign()先获取检索到的sourceDocuments
  • 基于sourceDocuments生成上下文字符串,供回答模板使用
  • 最后将模型生成的回答和原始来源文档合并成一个对象返回,这样就能同时拿到两者

内容的提问来源于stack exchange,提问作者Dante

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 00:07:35