AWS Lambda中Puppeteer下载CSV文件后无法定位的问题求助
解决AWS Lambda中Puppeteer下载CSV文件后无法定位并上传S3的问题
问题核心
在AWS Lambda环境下使用Puppeteer下载CSV文件,页面操作流程正常,但始终出现「找不到已下载文件」的错误,无法完成S3上传。本地环境测试正常,且Lambda已配置S3上传权限,问题出在文件下载路径和等待逻辑上。
错误原因分析
从CloudWatch日志和函数响应可明确:
- 代码中设置的下载目录指向了
/var/task/tmp(Lambda的代码部署目录),该目录为只读,Puppeteer无法写入文件,导致下载失败 - 使用固定
delay(8000)等待下载完成,无法保证文件完全写入磁盘,可能出现文件未生成就尝试读取的情况 - 硬编码文件名
Report.csv,可能与网站实际返回的文件名不符
解决方案
1. 强制使用Lambda可写临时目录
Lambda唯一支持读写的临时目录是/tmp,直接将下载路径设置为该目录,无需拼接代码目录路径:
const client = await page.createCDPSession(); const downloadDirectory = "/tmp"; // 直接使用Lambda的可写临时目录 await client.send("Page.setDownloadBehavior", { behavior: "allow", downloadPath: downloadDirectory, });
2. 监听下载事件,确保文件完全写入
替换固定delay,通过CDP会话监听下载开始和完成事件,动态获取文件名并等待下载结束:
// 监听下载事件,获取文件名并等待完成 let downloadedFileName; const downloadPromise = new Promise((resolve) => { client.on("Page.downloadWillBegin", (event) => { downloadedFileName = decodeURIComponent(path.basename(event.guid)); console.log("开始下载文件:", downloadedFileName); }); client.on("Page.downloadProgress", (event) => { if (event.state === "completed") { console.log("文件下载完成:", downloadedFileName); resolve(); } }); }); // 点击下载按钮 await page.click(".col-12.rightEnd .tableButton:nth-child(2)").catch(e => console.log(e)); // 等待下载完成 await downloadPromise;
3. 动态获取下载文件名,避免硬编码
使用监听事件中获取的downloadedFileName构建文件路径,确保和实际下载的文件名一致:
const filePath = path.join(downloadDirectory, downloadedFileName);
4. 可选:确保临时目录存在
虽然Lambda默认会创建/tmp目录,可添加代码确保目录存在,避免极端情况:
if (!fs.existsSync(downloadDirectory)) { fs.mkdirSync(downloadDirectory, { recursive: true }); }
完整修正后的关键代码片段
export async function handler(event) { const browser = await launch({ args: chromium.args, executablePath: await chromium.executablePath(), }); const page = (await browser.pages())[0]; await page.setDefaultNavigationTimeout(240000); const client = await page.createCDPSession(); const downloadDirectory = "/tmp"; // 确保目录存在 if (!fs.existsSync(downloadDirectory)) { fs.mkdirSync(downloadDirectory, { recursive: true }); } await client.send("Page.setDownloadBehavior", { behavior: "allow", downloadPath: downloadDirectory, }); // ... 省略页面登录、导航等操作代码 ... // 监听下载事件,等待完成 let downloadedFileName; const downloadPromise = new Promise((resolve) => { client.on("Page.downloadWillBegin", (event) => { downloadedFileName = decodeURIComponent(path.basename(event.guid)); console.log("开始下载:", downloadedFileName); }); client.on("Page.downloadProgress", (event) => { if (event.state === "completed") { console.log("下载完成:", downloadedFileName); resolve(); } }); }); // 触发下载 await page.click(".col-12.rightEnd .tableButton:nth-child(2)").catch(e => console.log(e)); await downloadPromise; // 构建正确的文件路径 const filePath = path.join(downloadDirectory, downloadedFileName); // 上传到S3 await uploadFileToS3(filePath); await browser.close(); return JSON.stringify({ message: "文件已上传至S3", filePath: filePath, fileName: downloadedFileName }); }
验证步骤
- 部署修正后的代码到Lambda
- 触发函数后查看CloudWatch日志,确认输出「开始下载」「下载完成」的日志
- 检查目标S3桶,确认文件已成功上传
内容的提问来源于stack exchange,提问作者Kelash
相关产品推荐
相关产品推荐

