在AWS Lambda上用Node.js+Langchain处理PPTX时遇ENOENT错误
AWS Lambda中Langchain处理PPTX文件报错:无法创建临时目录
我在AWS Lambda上开发Node.js应用,用于处理来自S3的PDF、CSV、TXT、JSON、DOCX及PPTX文件,拆分文本后存入Pinecone数据库,使用Langchain库处理文档加载与文本拆分。PDF、CSV、TXT、JSON、DOCX文件处理均正常,但处理PPTX文件时出现错误。
错误信息
2023-12-04T16:07:19.174Z 6d39f682-34ce-5441-af57-ab65cfa7facb ERROR Invoke Error { "errorType": "Error", "errorMessage": "[OfficeParser]: Error: ENOENT: no such file or directory, mkdir 'officeParserTemp/tempfiles'", "stack": [ "Error: [OfficeParser]: Error: ENOENT: no such file or directory, mkdir 'officeParserTemp/tempfiles'", " at Object.intoError (file:///var/runtime/index.mjs:46:16)", " at Object.textErrorLogger [as logError] (file:///var/runtime/index.mjs:684:56)", " at postError (file:///var/runtime/index.mjs:801:27)", " at done (file:///var/runtime/index.mjs:833:11)", " at fail (file:///var/runtime/index.mjs:845:11)", " at file:///var/runtime/index.mjs:872:20" ] }
相关代码片段
async function processFile(filename: string, key: string) { try { // Fetch the File content from S3 using the new method const command = new GetObjectCommand({ Bucket: process.env.S3_BUCKET_NAME, Key: filename, }); const content: any = await s3Client.send(command); const tempFilePath = path.join('/tmp', path.basename(filename)); await fs.writeFile(tempFilePath, content.Body); // Directly save the buffer // Load and split the File const fileExtension = path.extname(tempFilePath).toLowerCase(); let loader: any; switch (fileExtension) { case '.pdf': loader = new PDFLoader(tempFilePath); // working break; case '.csv': loader = new CSVLoader(tempFilePath); // working break; case '.txt': loader = new TextLoader(tempFilePath); //working break; case '.json': loader = new JSONLoader(tempFilePath); //working break; case '.docx': loader = new DocxLoader(tempFilePath); // working break; case '.pptx': console.log('LOADING THE PPTX FILE'); console.log('tempFilePath SRC', tempFilePath); loader = new PPTXLoader(tempFilePath); break; default: console.log('PROVIDED KEY: ', key); loader = new UnstructuredLoader(tempFilePath, { apiKey: key }); } const rawDocs = await loader.load(); const textSplitter = new RecursiveCharacterTextSplitter({ chunkSize: 1000, chunkOverlap: 200, }); const docs = await textSplitter.splitDocuments(rawDocs); // Generate embeddings and ingest into Pinecone const embeddings = new OpenAIEmbeddings(); const index = pinecone.index(process.env.PINECONE_INDEX_NAME); await PineconeStore.fromDocuments(docs, embeddings, { pineconeIndex: index, namespace: process.env.PINECONE_NAME_SPACE, textKey: 'text', }); } catch (error) { console.error(`Error processing and ingesting File: ${filename}. Error: ${error}`); throw error; } }
问题原因与解决方案
原因
AWS Lambda的文件系统仅/tmp目录具备可写权限,而Langchain的PPTXLoader依赖的OfficeParser库默认会在当前工作目录创建officeParserTemp/tempfiles临时目录,当前工作目录为只读,因此触发创建目录失败的错误。
解决方案
方案1:提前创建/tmp下的临时目录
在初始化PPTXLoader前,递归创建/tmp/officeParserTemp/tempfiles目录:
case '.pptx': console.log('LOADING THE PPTX FILE'); console.log('tempFilePath SRC', tempFilePath); // 递归创建可写的临时目录 const officeTempDir = path.join('/tmp', 'officeParserTemp', 'tempfiles'); fs.mkdirSync(officeTempDir, { recursive: true }); loader = new PPTXLoader(tempFilePath); break;
方案2:设置环境变量指定临时目录
通过设置环境变量让OfficeParser使用/tmp下的路径:
case '.pptx': console.log('LOADING THE PPTX FILE'); console.log('tempFilePath SRC', tempFilePath); // 设置OfficeParser的临时目录到/tmp process.env.OFFICE_PARSER_TMP_DIR = '/tmp/officeParserTemp'; loader = new PPTXLoader(tempFilePath); break;
方案3:使用PPTXLoader的临时目录配置(若支持)
检查Langchain的PPTXLoader是否支持传入临时目录参数,若支持可直接指定:
loader = new PPTXLoader(tempFilePath, { tempDir: '/tmp/officeParserTemp' });
内容的提问来源于stack exchange,提问作者Lloukas
相关产品推荐
相关产品推荐

