如何修改Puppeteer AWS Lambda代码批量处理30个URL截图任务
I've run into this exact issue before—spawning too many Chrome instances in quick succession via separate Lambda triggers eats up resources and causes browser crashes. Here's how to refactor your code to handle 30 URLs in a single batch, which will fix the disconnect error and improve efficiency:
Why Your Original Setup Failed
Your current approach triggers 30 independent Lambda functions every 2-3 seconds, each spinning up its own Chrome instance. Chrome is resource-intensive, and this rapid, parallel spawning overwhelms Lambda's execution environment resources, leading to browser processes crashing mid-navigation (the error you're seeing). Batching tasks lets you reuse one Chrome instance for all URLs, cutting down on overhead drastically.
Refactored Lambda Code
Replace your src/capture.js with this code, which handles batch payloads and reuses a single browser instance:
const chromeLambda = require("chrome-aws-lambda"); const S3Client = require("aws-sdk/clients/s3"); process.setMaxListeners(0); // Fix MaxListeners Error const s3 = new S3Client({ region: process.env.S3_REGION }); const defaultViewport = { width: 1920, height: 1080 }; exports.handler = async event => { let browser; const results = []; try { // Launch one browser to reuse for all tasks (key optimization!) browser = await chromeLambda.puppeteer.launch({ args: chromeLambda.args, executablePath: await chromeLambda.executablePath, defaultViewport, headless: chromeLambda.headless }); // Extract batch tasks from the incoming payload const tasks = Array.isArray(event.tasks) ? event.tasks : []; if (tasks.length === 0) { return { success: false, message: "No screenshot tasks provided" }; } // Process each task one at a time (parallel is possible but riskier with tab limits) for (const task of tasks) { const taskResult = { url: task.url, success: false }; try { console.log(`Starting screenshot for: ${task.url}`); // Open a new tab for each task const page = await browser.newPage(); // Navigate with timeout to avoid hanging on unresponsive sites await page.goto(task.url, { waitUntil: "networkidle2", timeout: 30000 }); // Capture screenshot const buffer = await page.screenshot(); // Use provided filename or auto-generate from domain const filename = task.filename || (new URL(task.url)).hostname.replace('www.', '') + '.png'; // Upload to S3 const s3Upload = await s3.upload({ Bucket: process.env.S3_BUCKET, Key: filename, Body: buffer, ContentType: "image/png", ACL: "public-read" }).promise(); taskResult.success = true; taskResult.s3Url = s3Upload.Location; console.log(`Completed screenshot for: ${task.url}`); } catch (taskErr) { taskResult.error = taskErr.message; console.error(`Failed to capture ${task.url}: ${taskErr.message}`); } finally { // Close the tab to free up memory (critical for batch processing) if (page) await page.close(); } results.push(taskResult); } } catch (browserErr) { console.error("Failed to launch Chrome:", browserErr.message); results.push({ success: false, error: "Browser initialization failed" }); } finally { // Always clean up the browser instance if (browser) await browser.close(); } // Return a summary of results return { totalTasks: results.length, successfulTasks: results.filter(r => r.success).length, taskResults: results }; };
Key Improvements in This Code
- Browser Reuse: Only one Chrome instance is launched for all 30 tasks, eliminating redundant resource-heavy initialization.
- Error Isolation: A failed screenshot for one URL won't break the entire batch—errors are logged per task.
- Resource Cleanup: Tabs are closed after each task, and the browser is always closed in the
finallyblock to prevent memory leaks. - Timeout Protection: The
page.gotocall includes a 30-second timeout to avoid hanging on unresponsive sites. - Flexible Filenames: You can either specify a filename in the payload or let the code generate one from the URL's domain.
Batch Payload Example
Send this single JSON payload to your API Gateway to trigger the batch processing:
{ "tasks": [ {"url": "https://gavurin.com", "filename": "gavurin.com.png"}, {"url": "https://example.com", "filename": "example.com.png"}, {"url": "https://github.com", "filename": "github.png"}, // Add 27 more URL/filename pairs here ] }
Additional Tips
- If you want faster execution, replace the
for...ofloop withPromise.allto process tasks in parallel. Note that Chrome has a soft limit on concurrent tabs (around 20-30), so this should work for your 30-task batch. - Monitor Lambda memory usage: Batch processing may require slightly more memory than single-task execution—start with 1024MB and adjust if you see performance issues.
内容的提问来源于stack exchange,提问作者sigur7

