You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何合并LangChain Documents并正确导入Pinecone向量库?

问题:合并PDF文档与视频转录文档并导入Pinecone向量库

我需要将PDF文件和视频转录文本导入Pinecone向量库,读取PDF和加载视频转录JSON的代码都能正常运行,但合并两类Document并导入时出现错误。

读取PDF的代码(正常运行)

const directoryLoaderPDF = new DirectoryLoader(filePath, {
  '.pdf': (path) => new CustomPDFLoader(path),
});

// const loader = new PDFLoader(filePath);
const rawPDFs = await directoryLoaderPDF.load();

/* Split text into chunks */
const textSplitter = new RecursiveCharacterTextSplitter({
  chunkSize: 1000,
  chunkOverlap: 200,
});

var docs = await textSplitter.splitDocuments(rawPDFs);

加载视频转录JSON的代码(正常运行)

const loaderJSON = new JSONLoader(
  'path',
);

const transcripts = await loaderJSON.load();

错误的合并方式及报错

当前用var documents: any[] = [docs, transcripts];合并,导入Pinecone时触发错误:

await PineconeStore.fromDocuments(documents, embeddings, {
  pineconeIndex: index,
  namespace: PINECONE_NAME_SPACE,
  textKey: 'text',
});

错误信息:

error [TypeError: Cannot read properties of undefined (reading 'replaceAll')]

问题原因

documents现在是二维数组(包含两个子数组,分别存储PDF文档和转录文档),但PineconeStore.fromDocuments要求传入一维的Document对象数组,二维数组会导致内部处理时无法正确读取文档的文本属性,从而抛出replaceAll相关的错误。


解决方案

将两个数组合并为一维数组,有两种常用方式:

方式1:使用扩展运算符(推荐)

// 合并两个一维数组为一个新的一维数组
const documents = [...docs, ...transcripts];

方式2:使用数组的concat方法

// concat返回合并后的新数组,不修改原数组
const documents = docs.concat(transcripts);

替换错误的合并代码后,再调用PineconeStore.fromDocuments就能正常导入向量库了。

内容的提问来源于stack exchange,提问作者Dawienchi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 21:06:02