基于LangChainJS与Pinecone的私有数据查询异常排查求助
问题:LangChain+Pinecone+OpenAI 私有数据查询失效
我想用LangChain给OpenAI GPT注入私有上下文,实现基于LLM的私有数据查询。当前实现逻辑是用LangChainJS加载指定路径文档,分割为文本块后存入Pinecone构建向量存储,再基于该存储创建问答链查询。但多数时候模型像没获取到私有数据上下文,偶尔能回答普通问题,甚至否认知晓简单内容。已尝试更换GPT模型、调整文本块大小、重建Pinecone存储,问题仍存在。
已知信息:Pinecone索引维度设为1536,其他维度会报错 Vector dimension 1536 does not match the dimension of the index 1000
原始实现代码
import { Document } from "langchain/document"; import { TextLoader } from "langchain/document_loaders/fs/text"; import { PDFLoader } from "langchain/document_loaders/fs/pdf"; import { CharacterTextSplitter } from "langchain/text_splitter"; import { PineconeClient } from "@pinecone-database/pinecone"; import { OpenAIEmbeddings } from "langchain/embeddings/openai"; import { PineconeStore } from "langchain/vectorstores/pinecone"; import { OpenAI } from "langchain/llms/openai"; import { VectorDBQAChain } from "langchain/chains"; const openAIApiKey = process.env.OPEN_AI_API_KEY; async function main(filePath) { // create document array const docs = [ new Document({ metadata: { name: `Filepath: ${filePath}` }, }), ]; // initialize loader const Loader = path.extname(file) === `.pdf` ? PDFLoader : TextLoader; const loader = new Loader(file); // load and split the docs const loadedAndSplitted = await loader.loadAndSplit(); // push the splitted docs to the array docs.push(...loadedAndSplitted); // create splitter const textSplitter = new CharacterTextSplitter({ chunkSize: 1000, chunkOverlap: 0, }); // use the splitter to split the docs to different chunks const splittedDocs = await textSplitter.splitDocuments(docs); // create pinecone index const client = new PineconeClient(); await client.init({ apiKey: process.env.PINECONE_API_KEY, environment: process.env.PINECONE_ENVIRONMENT, }); const pineconeIndex = client.Index(process.env.PINECONE_INDEX); // create openai embedding const embeddings = new OpenAIEmbeddings({ openAIApiKey }); // create a pinecone store using the splitted docs and the pinecone index const pineconeStore = await PineconeStore.fromDocuments( splittedDocs, embeddings, { pineconeIndex, namespace: "my-pinecode-index", } ); // initialize openai model const model = new OpenAI({ openAIApiKey, modelName: "gpt-3.5-turbo", }); // create a vector chain using the llm model and the pinecone store const chain = VectorDBQAChain.fromLLM(model, pineconeStore, { k: 1, returnSourceDocuments: true, }); // use the chain to query my data const response = await chain.call({ query: "Explain about the contents of the pdf file I provided.", // question is based on the file i provided }); console.log(` Response: ${response.text}`); }
排查与优化建议
1. 修复代码逻辑错误
- 变量名错误:代码中
path.extname(file)的file未定义,应改为传入的filePath,否则Loader初始化失败,根本没加载目标文档。 - 无效Document干扰:初始化的
docs数组里添加了一个无内容的空Document(仅含metadata),会被存入向量库,查询时可能匹配到该空文档,导致模型无有效上下文。 - 重复分割问题:先调用
loader.loadAndSplit()(默认用CharacterTextSplitter),之后又重复分割,导致文本块过度碎片化,破坏上下文连贯性。
2. 向量存储与查询优化
- 调整召回数量:当前
k:1仅召回1个文本块,若问题涉及内容分散在多个块中,会导致上下文缺失,建议改为k:3或k:5。 - 优化文本分割策略:当前
chunkOverlap:0会导致语义断裂,建议设置chunkOverlap:100-200;同时可根据文档类型调整chunkSize,比如PDF用800,纯文本用1200。 - 清理旧数据:每次运行往同一namespace写入数据,旧数据会干扰查询,建议写入前清理对应namespace的旧数据,或使用唯一namespace。
- 验证文档加载结果:在
loader.loadAndSplit()后打印内容,确认文档是否正确加载、是否有有效文本。
3. 模型与链的选择优化
- 模型适配:
OpenAI类针对文本补全模型(如davinci),gpt-3.5-turbo是聊天模型,应使用ChatOpenAI类(从langchain/chat_models/openai导入)。 - 更换问答链:
VectorDBQAChain已被标记为废弃,建议使用RetrievalQAChain,支持更多配置(如自定义prompt),能更好引导模型使用私有上下文。
修正后的代码示例
import { TextLoader } from "langchain/document_loaders/fs/text"; import { PDFLoader } from "langchain/document_loaders/fs/pdf"; import { CharacterTextSplitter } from "langchain/text_splitter"; import { PineconeClient } from "@pinecone-database/pinecone"; import { OpenAIEmbeddings } from "langchain/embeddings/openai"; import { PineconeStore } from "langchain/vectorstores/pinecone"; import { ChatOpenAI } from "langchain/chat_models/openai"; import { RetrievalQAChain } from "langchain/chains"; import path from "path"; // 导入path模块 const openAIApiKey = process.env.OPEN_AI_API_KEY; async function main(filePath) { // 初始化Loader,修复变量名错误 const ext = path.extname(filePath); const Loader = ext === ".pdf" ? PDFLoader : TextLoader; const loader = new Loader(filePath); // 配置文本分割器,一次性完成加载与分割 const textSplitter = new CharacterTextSplitter({ chunkSize: 800, chunkOverlap: 150, }); const loadedAndSplitted = await loader.loadAndSplit(textSplitter); // 初始化Pinecone const client = new PineconeClient(); await client.init({ apiKey: process.env.PINECONE_API_KEY, environment: process.env.PINECONE_ENVIRONMENT, }); const pineconeIndex = client.Index(process.env.PINECONE_INDEX); const embeddings = new OpenAIEmbeddings({ openAIApiKey }); // 清理旧数据(可选,按需使用) await pineconeIndex.delete1({ deleteAll: true, namespace: "my-pinecode-index", }); // 构建向量存储 const pineconeStore = await PineconeStore.fromDocuments( loadedAndSplitted, embeddings, { pineconeIndex, namespace: "my-pinecode-index", } ); // 使用ChatOpenAI适配gpt-3.5-turbo const model = new ChatOpenAI({ openAIApiKey, modelName: "gpt-3.5-turbo", temperature: 0, // 降低温度提升回答精准度 }); // 使用RetrievalQAChain替代废弃的VectorDBQAChain const chain = RetrievalQAChain.fromLLM(model, pineconeStore.asRetriever(3), { returnSourceDocuments: true, }); // 执行查询 const response = await chain.call({ query: "Explain about the contents of the pdf file I provided.", }); console.log(`Response: ${response.text}`); // 打印来源文档,验证是否匹配到正确内容 console.log("Source Documents:", response.sourceDocuments); }
内容的提问来源于stack exchange,提问作者javascript-wtf
相关产品推荐
相关产品推荐

