如何为Node.js Puppeteer爬取脚本添加下一页分页遍历逻辑
分页爬取实现方案
你可以将单页爬取逻辑封装为复用函数,通过循环检测「下一页」按钮的可点击状态实现全分页遍历,同时修复原有代码中waitForSelector未加await的潜在加载时序问题,完整修改后代码如下:
const puppeteer = require("puppeteer"); const base = "https://www.timesbusinessdirectory.com"; const url = "https://www.timesbusinessdirectory.com/company-listings"; // 封装单页爬取逻辑:爬取当前列表页所有企业详情数据 async function scrapeCurrentPage(page) { await page.waitForSelector(".company-listing"); const sections = await page.$$(".company-listing"); const linkList = []; // 先收集当前页所有企业内页链接 for (const section of sections) { const itemHref = await section.$eval("h3 > a", el => el.getAttribute("href")); linkList.push(itemHref); } // 遍历爬取每个企业详情 for (const link of linkList) { const detailUrl = base + link; await page.goto(detailUrl, {waitUntil: 'networkidle2'}); await page.waitForSelector(".company-details"); const detailBlocks = await page.$$(".company-details"); for (const block of detailBlocks) { const company = await block.$eval("h3", el => el.textContent.trim()); const phone = await block.$eval("#valuephone a[href]", el => el.textContent.trim()); const emailOnclick = await block.$eval("a[onclick^='showCompanyEmail']", el => el.getAttribute("onclick")); const email = emailOnclick.split("('")[1].split("',")[0]; console.log({title: company, tel: phone, emailId: email}); } } } (async () => { const browser = await puppeteer.launch({headless: false}); const [page] = await browser.pages(); // 打开初始列表页 await page.goto(url, {waitUntil: 'networkidle2'}); while (true) { // 爬取当前页所有数据 await scrapeCurrentPage(page); // 返回列表页 await page.goBack({waitUntil: 'networkidle2'}); // 检测下一页按钮是否可点击 const nextBtn = await page.$('.pagination li:last-child a'); const isLastPage = await page.$eval('.pagination li:last-child', el => el.classList.contains('disabled')); if (isLastPage || !nextBtn) break; // 点击下一页,等待页面加载完成 await Promise.all([ nextBtn.click(), page.waitForNavigation({waitUntil: 'networkidle2'}) ]); } await browser.close(); })();
补充说明
- 可根据需求将
headless参数改为true实现后台无界面运行 - 若担心请求频率过高被站点限制,可在页面跳转逻辑后添加
await page.waitForTimeout(1000)类的延迟逻辑 - 若爬取过程中出现偶现的元素找不到报错,可调整
waitForSelector的超时参数,或者添加重试逻辑
内容的提问来源于stack exchange,提问作者robots.txt
相关产品推荐
相关产品推荐

