如何让Promise.all执行完毕后再输出结果?Node.js Puppeteer问题
问题:Puppeteer多网页爬取Promise执行顺序异常
我正在使用Node.js的Puppeteer进行多网页爬取,每个爬取流程对应一个Promise,希望所有Promise执行完成后再输出"all finished"。但当前代码输出顺序为:先打印"all finished",随后才是三次"partial end",请问如何调整让"all finished"在最后输出?
当前输出
all finished partial end partial end partial end
原代码
const puppeteer = require("puppeteer"); const URLs = [ 'https://www.google.es/search?q=dog&sxsrf=ALiCzsaZ5RIpFrQHMAxy9uZ9vbCu2wDAlw:1662240805949&source=lnms&tbm=isch&sa=X&ved=2ahUKEwjp7ZPGyfn5AhVEgv0HHSbDC1oQ_AUoAXoECAIQAw&biw=1280&bih=576&dpr=2', 'https://www.google.es/search?q=dog&sxsrf=ALiCzsaZ5RIpFrQHMAxy9uZ9vbCu2wDAlw:1662240805949&source=lnms&tbm=isch&sa=X&ved=2ahUKEwjp7ZPGyfn5AhVEgv0HHSbDC1oQ_AUoAXoECAIQAw&biw=1280&bih=576&dpr=2', 'https://www.google.es/search?q=dog&sxsrf=ALiCzsaZ5RIpFrQHMAxy9uZ9vbCu2wDAlw:1662240805949&source=lnms&tbm=isch&sa=X&ved=2ahUKEwjp7ZPGyfn5AhVEgv0HHSbDC1oQ_AUoAXoECAIQAw&biw=1280&bih=576&dpr=2' ]; main(); async function main() { await scrapingUrls().then(console.log("all finished")); } async function scrapingUrls() { const prom = []; for (var i = 0; i < URLs.length; i++) { prom.push(scrapInformation(URLs[i])); } return await Promise.all([prom]); } function scrapInformation(url) { return new Promise(async function(resolve, reject) { const browser = await puppeteer.launch() const page = await browser.newPage() await page.goto(url, {waitUntil: 'networkidle2'}); await browser.close().then(function () { console.log('partial end'); resolve(); }) }); }
解决方案
问题出在三个关键细节上,修改后即可让"all finished"最后输出:
1. 修复main函数的then回调写法
原代码中then(console.log("all finished"))是立即执行console.log,而非等待scrapingUrls完成后执行。改成以下两种写法都可以:
async function main() { await scrapingUrls(); console.log("all finished"); // 或者用then回调: // await scrapingUrls().then(() => console.log("all finished")); }
2. 修复scrapingUrls中的Promise.all调用
Promise.all接受的是Promise数组,原代码里[prom]把数组又嵌套了一层,导致Promise.all等待的是一个包含数组的Promise,而非直接等待所有爬取任务。应该直接传入prom数组:
async function scrapingUrls() { const prom = []; for (var i = 0; i < URLs.length; i++) { prom.push(scrapInformation(URLs[i])); } return await Promise.all(prom); // 去掉外层[] }
3. 优化scrapInformation的Promise写法
原代码用new Promise包裹async函数属于冗余写法,直接把scrapInformation改成async函数更简洁,同时调整browser.close后的逻辑:
async function scrapInformation(url) { const browser = await puppeteer.launch(); const page = await browser.newPage(); await page.goto(url, {waitUntil: 'networkidle2'}); await browser.close(); console.log('partial end'); }
修改后的完整代码
const puppeteer = require("puppeteer"); const URLs = [ 'https://www.google.es/search?q=dog&sxsrf=ALiCzsaZ5RIpFrQHMAxy9uZ9vbCu2wDAlw:1662240805949&source=lnms&tbm=isch&sa=X&ved=2ahUKEwjp7ZPGyfn5AhVEgv0HHSbDC1oQ_AUoAXoECAIQAw&biw=1280&bih=576&dpr=2', 'https://www.google.es/search?q=dog&sxsrf=ALiCzsaZ5RIpFrQHMAxy9uZ9vbCu2wDAlw:1662240805949&source=lnms&tbm=isch&sa=X&ved=2ahUKEwjp7ZPGyfn5AhVEgv0HHSbDC1oQ_AUoAXoECAIQAw&biw=1280&bih=576&dpr=2', 'https://www.google.es/search?q=dog&sxsrf=ALiCzsaZ5RIpFrQHMAxy9uZ9vbCu2wDAlw:1662240805949&source=lnms&tbm=isch&sa=X&ved=2ahUKEwjp7ZPGyfn5AhVEgv0HHSbDC1oQ_AUoAXoECAIQAw&biw=1280&bih=576&dpr=2' ]; main(); async function main() { await scrapingUrls(); console.log("all finished"); } async function scrapingUrls() { const prom = []; for (var i = 0; i < URLs.length; i++) { prom.push(scrapInformation(URLs[i])); } return await Promise.all(prom); } async function scrapInformation(url) { const browser = await puppeteer.launch(); const page = await browser.newPage(); await page.goto(url, {waitUntil: 'networkidle2'}); await browser.close(); console.log('partial end'); }
修改后输出顺序会变为:
partial end partial end partial end all finished
内容的提问来源于stack exchange,提问作者solamente
相关产品推荐
相关产品推荐

