使用Puppeteer爬取8000+子页面获取商品EAM ID的技术问询
Puppeteer爬取电商EAM ID的异步遍历优化与实践
嘿,我看了你用Puppeteer爬取电商EAM ID的代码,先给你点个赞——用asyncForEach来处理异步遍历的思路是对的,不过咱们可以把它优化得更顺手,同时结合Puppeteer的场景给你一些实践建议:
1. 优化你的asyncForEach函数
你当前实现的asyncForEach功能没问题,但可以用ES6模板字符串让日志更简洁,甚至用for...of循环简化代码结构:
简化后的版本
async function asyncForEach(array, callback) { let index = 0; for (const item of array) { console.log(`Iterating through array ${index} Of ${array.length}`); await callback(item, index++, array); } }
这个版本和你原来的逻辑完全一致,只是写法更现代、可读性更强。
2. 结合Puppeteer的EAM ID爬取实现
假设productsToview是产品页面的URL列表,在遍历回调里可以这样标准化地获取EAM ID,同时做好错误处理和资源释放:
// 假设你已经初始化了browser实例 await asyncForEach(productsToview, async (productUrl, index) => { let page; try { // 打开新页面,也可以考虑复用页面提升效率 page = await browser.newPage(); // 等待页面加载完成(networkidle2适合电商页面,减少等待时间) await page.goto(productUrl, { waitUntil: 'networkidle2', timeout: 60000 }); // 根据页面实际结构获取EAM ID,以下是两种常见场景: // 场景1:EAM ID在DOM元素中,用选择器获取 const eamId = await page.$eval('.product-eam-id', el => el.textContent.trim()); // 场景2:EAM ID存在于页面全局变量中 // const eamId = await page.evaluate(() => window.productDetails.eamId); console.log(`✅ 第${index+1}个商品EAM ID: ${eamId}`); // 这里可以把eamId存入数组、写入文件或者数据库 } catch (error) { console.error(`❌ 处理第${index+1}个商品失败: ${error.message}`); } finally { // 无论成功失败,都关闭页面释放浏览器资源 if (page) await page.close(); } });
3. 可选:并发处理提升爬取效率
如果你的商品列表很长,串行遍历速度太慢,可以限制并发数(比如同时处理3个页面),避免浏览器因资源耗尽崩溃:
const concurrencyLimit = 3; // 控制同时打开的页面数 let activeTasks = 0; const taskQueue = []; async function processProduct(productUrl, index) { activeTasks++; let page; try { page = await browser.newPage(); await page.goto(productUrl, { waitUntil: 'networkidle2', timeout: 60000 }); const eamId = await page.$eval('.product-eam-id', el => el.textContent.trim()); console.log(`✅ 第${index+1}个商品EAM ID: ${eamId}`); } catch (error) { console.error(`❌ 处理第${index+1}个商品失败: ${error.message}`); } finally { if (page) await page.close(); activeTasks--; // 队列里有任务就继续执行 if (taskQueue.length > 0) { taskQueue.shift()(); } } } // 遍历商品列表,控制并发 for (let index = 0; index < productsToview.length; index++) { const productUrl = productsToview[index]; if (activeTasks < concurrencyLimit) { processProduct(productUrl, index); } else { taskQueue.push(() => processProduct(productUrl, index)); } }
这样既保证了爬取效率,又不会给浏览器带来太大压力。
内容的提问来源于stack exchange,提问作者Marc Rasmussen
相关产品推荐
相关产品推荐

