You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于LangChainJS与Pinecone的私有数据查询异常排查求助

问题:LangChain+Pinecone+OpenAI 私有数据查询失效

我想用LangChain给OpenAI GPT注入私有上下文,实现基于LLM的私有数据查询。当前实现逻辑是用LangChainJS加载指定路径文档,分割为文本块后存入Pinecone构建向量存储,再基于该存储创建问答链查询。但多数时候模型像没获取到私有数据上下文,偶尔能回答普通问题,甚至否认知晓简单内容。已尝试更换GPT模型、调整文本块大小、重建Pinecone存储,问题仍存在。

已知信息:Pinecone索引维度设为1536,其他维度会报错 Vector dimension 1536 does not match the dimension of the index 1000


原始实现代码

import { Document } from "langchain/document";
import { TextLoader } from "langchain/document_loaders/fs/text";
import { PDFLoader } from "langchain/document_loaders/fs/pdf";
import { CharacterTextSplitter } from "langchain/text_splitter";

import { PineconeClient } from "@pinecone-database/pinecone";

import { OpenAIEmbeddings } from "langchain/embeddings/openai";
import { PineconeStore } from "langchain/vectorstores/pinecone";

import { OpenAI } from "langchain/llms/openai";
import { VectorDBQAChain } from "langchain/chains";

const openAIApiKey = process.env.OPEN_AI_API_KEY;

async function main(filePath) {
  // create document array
  const docs = [
    new Document({
      metadata: { name: `Filepath: ${filePath}` },
    }),
  ];

  // initialize loader
  const Loader = path.extname(file) === `.pdf` ? PDFLoader : TextLoader;

  const loader = new Loader(file);

  // load and split the docs
  const loadedAndSplitted = await loader.loadAndSplit();

  // push the splitted docs to the array
  docs.push(...loadedAndSplitted);

  // create splitter
  const textSplitter = new CharacterTextSplitter({
    chunkSize: 1000,
    chunkOverlap: 0,
  });

  // use the splitter to split the docs to different chunks
  const splittedDocs = await textSplitter.splitDocuments(docs);

  // create pinecone index
  const client = new PineconeClient();
  await client.init({
    apiKey: process.env.PINECONE_API_KEY,
    environment: process.env.PINECONE_ENVIRONMENT,
  });
  const pineconeIndex = client.Index(process.env.PINECONE_INDEX);

  // create openai embedding
  const embeddings = new OpenAIEmbeddings({ openAIApiKey });

  // create a pinecone store using the splitted docs and the pinecone index
  const pineconeStore = await PineconeStore.fromDocuments(
    splittedDocs,
    embeddings,
    {
      pineconeIndex,
      namespace: "my-pinecode-index",
    }
  );

  // initialize openai model
  const model = new OpenAI({
    openAIApiKey,
    modelName: "gpt-3.5-turbo",
  });

  // create a vector chain using the llm model and the pinecone store
  const chain = VectorDBQAChain.fromLLM(model, pineconeStore, {
    k: 1,
    returnSourceDocuments: true,
  });

  // use the chain to query my data
  const response = await chain.call({
    query: "Explain about the contents of the pdf file I provided.", // question is based on the file i provided
  });

  console.log(`
Response: ${response.text}`); 
}

排查与优化建议

1. 修复代码逻辑错误

  • 变量名错误:代码中path.extname(file)的file未定义,应改为传入的filePath,否则Loader初始化失败,根本没加载目标文档。
  • 无效Document干扰:初始化的docs数组里添加了一个无内容的空Document(仅含metadata),会被存入向量库,查询时可能匹配到该空文档,导致模型无有效上下文。
  • 重复分割问题:先调用loader.loadAndSplit()(默认用CharacterTextSplitter),之后又重复分割,导致文本块过度碎片化,破坏上下文连贯性。

2. 向量存储与查询优化

  • 调整召回数量:当前k:1仅召回1个文本块,若问题涉及内容分散在多个块中,会导致上下文缺失,建议改为k:3或k:5。
  • 优化文本分割策略:当前chunkOverlap:0会导致语义断裂,建议设置chunkOverlap:100-200;同时可根据文档类型调整chunkSize,比如PDF用800,纯文本用1200。
  • 清理旧数据:每次运行往同一namespace写入数据,旧数据会干扰查询,建议写入前清理对应namespace的旧数据,或使用唯一namespace。
  • 验证文档加载结果:在loader.loadAndSplit()后打印内容,确认文档是否正确加载、是否有有效文本。

3. 模型与链的选择优化

  • 模型适配:OpenAI类针对文本补全模型(如davinci),gpt-3.5-turbo是聊天模型,应使用ChatOpenAI类(从langchain/chat_models/openai导入)。
  • 更换问答链:VectorDBQAChain已被标记为废弃,建议使用RetrievalQAChain,支持更多配置(如自定义prompt),能更好引导模型使用私有上下文。

修正后的代码示例

import { TextLoader } from "langchain/document_loaders/fs/text";
import { PDFLoader } from "langchain/document_loaders/fs/pdf";
import { CharacterTextSplitter } from "langchain/text_splitter";
import { PineconeClient } from "@pinecone-database/pinecone";
import { OpenAIEmbeddings } from "langchain/embeddings/openai";
import { PineconeStore } from "langchain/vectorstores/pinecone";
import { ChatOpenAI } from "langchain/chat_models/openai";
import { RetrievalQAChain } from "langchain/chains";
import path from "path"; // 导入path模块

const openAIApiKey = process.env.OPEN_AI_API_KEY;

async function main(filePath) {
  // 初始化Loader,修复变量名错误
  const ext = path.extname(filePath);
  const Loader = ext === ".pdf" ? PDFLoader : TextLoader;
  const loader = new Loader(filePath);

  // 配置文本分割器,一次性完成加载与分割
  const textSplitter = new CharacterTextSplitter({
    chunkSize: 800,
    chunkOverlap: 150,
  });
  const loadedAndSplitted = await loader.loadAndSplit(textSplitter);

  // 初始化Pinecone
  const client = new PineconeClient();
  await client.init({
    apiKey: process.env.PINECONE_API_KEY,
    environment: process.env.PINECONE_ENVIRONMENT,
  });
  const pineconeIndex = client.Index(process.env.PINECONE_INDEX);
  const embeddings = new OpenAIEmbeddings({ openAIApiKey });

  // 清理旧数据(可选,按需使用)
  await pineconeIndex.delete1({
    deleteAll: true,
    namespace: "my-pinecode-index",
  });

  // 构建向量存储
  const pineconeStore = await PineconeStore.fromDocuments(
    loadedAndSplitted,
    embeddings,
    {
      pineconeIndex,
      namespace: "my-pinecode-index",
    }
  );

  // 使用ChatOpenAI适配gpt-3.5-turbo
  const model = new ChatOpenAI({
    openAIApiKey,
    modelName: "gpt-3.5-turbo",
    temperature: 0, // 降低温度提升回答精准度
  });

  // 使用RetrievalQAChain替代废弃的VectorDBQAChain
  const chain = RetrievalQAChain.fromLLM(model, pineconeStore.asRetriever(3), {
    returnSourceDocuments: true,
  });

  // 执行查询
  const response = await chain.call({
    query: "Explain about the contents of the pdf file I provided.",
  });

  console.log(`Response: ${response.text}`);
  // 打印来源文档,验证是否匹配到正确内容
  console.log("Source Documents:", response.sourceDocuments);
}

内容的提问来源于stack exchange,提问作者javascript-wtf

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 11:30:24