You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

LlamaIndex.ts管道部分Chunk Embeddings为undefined问题求助

LlamaIndex.ts IngestionPipeline处理长文档时多数Chunk嵌入为undefined且创建索引报错

问题现象

使用LlamaIndex.ts的IngestionPipeline处理文档,按官方文档实现OpenAIEmbedding后,文档分割仅前2个Chunk能正常生成嵌入向量,其余Chunk的embeddings均为undefined,创建VectorStoreIndex时报错:TypeError: Cannot read properties of undefined (reading 'filename')。调整OpenAIEmbedding的maxRetries、timeout、embedBatchSize等参数无效,仅文本较短(分割后≤2个Chunk)时能正常运行。

报错信息

this are the index: " Promise { <pending> }
Error indexing text:  TypeError: Cannot read properties of undefined (reading 'filename')

相关代码

const { Document, VectorStoreIndex, IngestionPipeline, SimpleNodeParser, OpenAIEmbedding } = require("llamaindex");

const processDocuments = async (text) => {
    try{
        const document = new Document({ text: text });
        const pipeline = new IngestionPipeline({
            transformations: [
                new SimpleNodeParser({ chunkSize: 1024, chunkOverlap: 20 }),
                new OpenAIEmbedding(),
            ],
        });
        const nodes = await pipeline.run({ documents: [document] }); 
        const index = VectorStoreIndex.fromDocuments(nodes.documents);
        console.log('this are the index: "', index)
    } catch (error) {
        console.log('Something wrong while processing the document: ', error);
        throw error;
    }
};
    
module.exports = {processDocuments}

解决方案

1. 纠正VectorStoreIndex的调用方式

IngestionPipeline.run()返回的是处理后的Node数组,而非包含documents属性的对象。你错误地使用了nodes.documents,导致传入fromDocuments的参数为undefined,这是触发filename读取错误的直接原因。

应改用VectorStoreIndex.fromNodes()方法,直接传入pipeline输出的nodes数组,同时给方法添加await避免输出pending状态的Promise:

// 替换原索引创建代码
const index = await VectorStoreIndex.fromNodes(nodes);
console.log('this are the index: "', index)

2. 排查OpenAIEmbedding批量嵌入失败原因

多数Chunk嵌入为undefined,可能是OpenAI API批量请求时出现静默失败。可以通过以下方式排查:

  • 在pipeline中添加自定义日志中间件,打印每个Node的处理状态,确认是否有异常文本导致嵌入失败:
const pipeline = new IngestionPipeline({
    transformations: [
        new SimpleNodeParser({ chunkSize: 1024, chunkOverlap: 20 }),
        // 添加日志中间件
        (nodes) => {
            console.log(`Processing ${nodes.length} nodes`);
            nodes.forEach((node, idx) => console.log(`Node ${idx} text length: ${node.text.length}`));
            return nodes;
        },
        new OpenAIEmbedding(),
    ],
});
  • 显式设置OpenAIEmbedding的embedBatchSize为较小值(比如10),避免单次请求过大触发API限制:
new OpenAIEmbedding({ embedBatchSize: 10 })

3. 确保OpenAI API配置正确

检查是否正确设置了OpenAI API密钥,可通过环境变量OPENAI_API_KEY或在初始化时显式传入:

new OpenAIEmbedding({ apiKey: "你的API密钥" })

内容的提问来源于stack exchange,提问作者lishing

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 21:55:10