You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在AWS Lambda中使用Puppeteer下载文件并上传至S3

在AWS Lambda中使用Puppeteer下载文件并上传至S3的实现指南

核心结论

完全可以在AWS Lambda中实现该需求,核心依赖确实是puppeteer-core(轻量版,不带Chromium)和@sparticuz/chromium(适配Lambda环境的无头Chromium),以下是详细适配步骤及代码改造方案:

步骤1:依赖配置

在项目package.json中添加以下依赖:

{
  "dependencies": {
    "puppeteer-core": "^21.0.0",
    "@sparticuz/chromium": "^116.0.0",
    "@aws-sdk/client-s3": "^3.400.0"
  }
}

注意:不要使用完整的puppeteer包,它自带的Chromium无法适配Lambda的Linux环境;@sparticuz/chromium是专门为Lambda优化的Chromium版本,体积更小且兼容Lambda运行时。

步骤2:Lambda环境配置

  • 内存配置:至少设置为1GB(推荐1.5GB),Chromium运行需要足够内存,内存不足会导致崩溃。
  • 超时设置:根据页面加载和下载耗时,设置为3-5分钟(Lambda最大支持15分钟)。
  • IAM权限:给Lambda执行角色添加S3读写权限(可配置精细权限,仅允许操作目标存储桶),确保能上传文件到S3。

步骤3:代码改造要点

3.1 Chromium启动配置

替换本地的puppeteer.launch配置,适配Lambda无沙箱环境:

const chromium = require('@sparticuz/chromium');
const puppeteer = require('puppeteer-core');

const browser = await puppeteer.launch({
  args: chromium.args,
  defaultViewport: chromium.defaultViewport,
  executablePath: await chromium.executablePath(),
  headless: chromium.headless,
  ignoreHTTPSErrors: true,
});

3.2 下载路径与完成监听

Lambda仅/tmp目录可写,修改下载路径为/tmp;同时需要监听下载完成事件,避免代码提前执行后续逻辑:

// 设置下载行为
await client.send('Page.setDownloadBehavior', {
  behavior: 'allow',
  downloadPath: '/tmp'
});

// 监听下载完成事件
page.on('download', async (download) => {
  const filePath = `/tmp/${download.suggestedFilename()}`;
  await download.saveAs(filePath);
  // 可在此触发S3上传逻辑
});

3.3 S3上传实现

使用AWS SDK v3的S3客户端完成文件上传:

const { S3Client, PutObjectCommand } = require('@aws-sdk/client-s3');

const s3Client = new S3Client({ region: '你的S3桶区域' });

async function uploadToS3(filePath, bucketName, key) {
  const fileContent = fs.readFileSync(filePath);
  const command = new PutObjectCommand({
    Bucket: bucketName,
    Key: key,
    Body: fileContent
  });
  await s3Client.send(command);
}

完整适配后的Lambda代码

const puppeteer = require('puppeteer-core');
const chromium = require('@sparticuz/chromium');
const fs = require('fs');
const { S3Client, PutObjectCommand } = require('@aws-sdk/client-s3');

// 初始化S3客户端
const s3Client = new S3Client({ region: 'us-east-1' }); // 替换为你的区域
const S3_BUCKET_NAME = 'your-bucket-name'; // 替换为你的S3桶名

async function uploadToS3(filePath, key) {
  try {
    const fileContent = fs.readFileSync(filePath);
    const command = new PutObjectCommand({
      Bucket: S3_BUCKET_NAME,
      Key: key,
      Body: fileContent
    });
    await s3Client.send(command);
    console.log('文件已成功上传至S3');
  } catch (err) {
    console.error('S3上传失败', err);
    throw err;
  }
}

exports.handler = async (event) => {
  let browser;
  try {
    // 启动适配Lambda的Chromium
    browser = await puppeteer.launch({
      args: chromium.args,
      defaultViewport: chromium.defaultViewport,
      executablePath: await chromium.executablePath(),
      headless: chromium.headless,
      ignoreHTTPSErrors: true,
    });

    const page = await browser.newPage();
    page.setDefaultNavigationTimeout(2 * 60 * 1000);

    // 设置下载行为到/tmp目录
    const client = await page.target().createCDPSession();
    await client.send('Page.setDownloadBehavior', {
      behavior: 'allow',
      downloadPath: '/tmp'
    });

    // 登录流程(建议通过Lambda环境变量传递敏感信息)
    await page.goto('https://app.website.com/login');
    await page.type('#email', process.env.EMAIL);
    await page.type('#password', process.env.PASSWORD);
    await page.click('[data-cy="submit"]');
    await page.waitForNavigation();

    // 触发下载并监听完成
    const el = await page.$x('//*[@id="download_button"]/div/a[1]');
    if (el.length > 0) {
      const downloadPromise = new Promise((resolve) => {
        page.on('download', async (download) => {
          const fileName = download.suggestedFilename();
          const filePath = `/tmp/${fileName}`;
          await download.saveAs(filePath);
          resolve(filePath);
        });
      });

      await el[0].click();
      const downloadedFilePath = await downloadPromise;

      // 上传到S3,添加时间戳避免文件名重复
      await uploadToS3(downloadedFilePath, `downloads/${Date.now()}-${downloadedFilePath.split('/').pop()}`);
    } else {
      console.log('未找到下载按钮');
      throw new Error('Download button not found');
    }

    return { statusCode: 200, body: '文件下载并上传S3成功' };
  } catch (e) {
    console.error('执行失败', e);
    return { statusCode: 500, body: JSON.stringify({ error: e.message }) };
  } finally {
    if (browser) {
      await browser.close();
    }
  }
};

部署注意事项

  1. 打包方式:本地打包时需针对Linux环境安装依赖,避免Mac/Windows二进制文件不兼容:
    npm install --platform=linux --arch=x64 --no-save
    
    随后将代码与node_modules一起打包为zip上传至Lambda,或使用SAM/Serverless Framework部署。
  2. Lambda层复用:可将@sparticuz/chromium和puppeteer-core打包成Lambda层,减少部署包体积并方便复用。
  3. 敏感信息处理:切勿硬编码账号密码,使用Lambda环境变量存储,通过process.env读取。

内容的提问来源于stack exchange,提问作者kajl16

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 02:35:01