You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Node.js版AWS Lambda无法提取PDF/文档文本的问题求助

问题:AWS Lambda部署后无法提取PDF文本(本地运行正常)

我用Node.js编写了AWS Lambda函数,目标是提取PDF、doc、docx文件中的文本并转换为单词数组。试了多个npm包,发现提取函数始终未执行。以下是使用pdf.js-extract的代码示例,本地运行完全正常,但部署到AWS Lambda后无法提取任何文本。

import { PDFExtract } from "pdf.js-extract";

...

export const handler = async (event) => {

    ...

    const pdfExtract = new PDFExtract();
    const tempFilePath = join(tmpdir(), "resume.pdf");
    const buffer = readFileSync(tempFilePath);
    const wordsList = [];

    await pdfExtract.extractBuffer(buffer, {}, (err, data) => {
        if (err)
            return console.log(err);
        data.pages[0].content.forEach((e) => {
            const str = e.str.trim().split(" ");
            str.forEach((word) => {
                if (word.length > 1)
                    wordsList.push(word);
            });
        });
    });
        console.log(wordsList);
}

核心问题分析

你这里的关键问题是混用了回调函数和async/await。pdf.js-extract的extractBuffer是基于回调的API,await对它不起作用——Lambda函数会在回调执行前就提前结束,导致wordsList还没被填充就输出空数组。

另外,Lambda运行环境和本地环境存在差异,还有几个细节需要注意。

修复步骤

1. 将回调转为Promise(关键修复)

把extractBuffer包装成Promise,让await能正确等待提取操作完成:

import { PDFExtract } from "pdf.js-extract";
import { join } from "path";
import { tmpdir, readFileSync } from "fs";

export const handler = async (event) => {
    const pdfExtract = new PDFExtract();
    const tempFilePath = join(tmpdir(), "resume.pdf");
    const buffer = readFileSync(tempFilePath);
    const wordsList = [];

    // 把回调式API包装成Promise
    const extractPDF = () => {
        return new Promise((resolve, reject) => {
            pdfExtract.extractBuffer(buffer, {}, (err, data) => {
                if (err) reject(err);
                else resolve(data);
            });
        });
    };

    try {
        const data = await extractPDF();
        // 处理提取到的文本
        data.pages[0].content.forEach((e) => {
            const str = e.str.trim().split(" ");
            str.forEach((word) => {
                if (word.length > 1)
                    wordsList.push(word);
            });
        });
        console.log(wordsList);
    } catch (err) {
        console.error("PDF提取失败:", err);
    }
};

2. 适配Lambda部署环境

  • 依赖包适配Linux环境:Lambda运行在Linux系统上,本地如果是Windows/Mac,打包前用npm install --platform=linux --arch=x64安装依赖,避免因系统差异导致依赖无法运行。
  • 调整内存配置:PDF提取需要一定内存,建议将Lambda内存配置至少设为512MB,防止内存不足中断提取。
  • 确认临时文件存在:Lambda的tmpdir()对应/tmp路径,要确保resume.pdf已被正确写入该路径(比如从S3下载后写入)。

3. doc/docx文件处理建议

如果要处理doc/docx格式,推荐使用mammoth包,它支持Promise风格API,更适配async/await:

import mammoth from "mammoth";

// 提取docx文本并转为单词数组的示例
const extractDocxWords = async (buffer) => {
    const result = await mammoth.extractRawText({ buffer });
    return result.value.split(" ").filter(word => word.length > 1);
};

内容的提问来源于stack exchange,提问作者Yinon Ozery

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 23:30:23