You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Node.js Puppeteer网页爬虫翻页问题:异步await使用及按钮点击困惑

Hey there! Let's work through your Puppeteer crawler issues step by step—first fixing that await is only valid in async functions error, then tackling the pagination click logic.


First: Fix the await Error

The root cause here is straightforward: any line using await must live inside an async function. Since Puppeteer’s core methods (like launch(), goto(), evaluate()) all return Promises, you need to wrap your entire crawler logic in an async context.

Here’s the basic structure you should use (this eliminates the error immediately):

const puppeteer = require('puppeteer');

// Wrap everything in an async IIFE (Immediately Invoked Function Expression)
(async () => {
  // Your crawler code goes here—all await calls will work now!
  const browser = await puppeteer.launch();
  const page = await browser.newPage();
  await page.goto('your-target-url');

  // ... rest of your logic ...

  await browser.close();
})();

This is the standard pattern for Puppeteer scripts because it lets you use await without having to handle raw Promise .then() chains.


Second: Implement Reliable Pagination

Now for the tricky part: detecting and clicking the "next page" button repeatedly until there are no more pages. Let’s break this into a robust loop with safeguards:

Key Steps for Pagination:

  1. Initialize an array to store all scraped data
  2. Loop indefinitely until no "next page" button exists
  3. Scrape the current page’s data and add it to your array
  4. Check if the next page button is present
  5. If it exists: click it, wait for the page to load (handle both full page reloads and AJAX loads)
  6. If it doesn’t exist: exit the loop

Full Working Example

Replace the selectors (.item-class, .next-page-button) with ones that match your target website:

const puppeteer = require('puppeteer');

(async () => {
  // Launch browser with headless: false for easy debugging
  const browser = await puppeteer.launch({ headless: false });
  const page = await browser.newPage();
  
  // Navigate to the target page, wait for network to settle
  await page.goto('your-target-url', { waitUntil: 'networkidle2' });

  const allScrapedData = [];

  while (true) {
    // Step 1: Scrape current page data (run code in browser context)
    const currentPageData = await page.evaluate(() => {
      // Replace this with your actual data extraction logic
      const items = document.querySelectorAll('.item-class');
      return Array.from(items).map(item => ({
        title: item.querySelector('.title').textContent.trim(),
        price: item.querySelector('.price').textContent.trim()
        // Add more fields as needed
      }));
    });

    // Merge current page data into the main array
    allScrapedData.push(...currentPageData);
    console.log(`Scraped ${currentPageData.length} items from current page`);

    // Step 2: Check for next page button
    const nextPageButton = await page.$('.next-page-button');
    if (!nextPageButton) {
      console.log('No more pages to scrape!');
      break; // Exit loop when no next button exists
    }

    // Step 3: Click next page and wait for content to load
    try {
      // Handle full page reloads: wait for navigation to complete
      await Promise.all([
        nextPageButton.click(),
        page.waitForNavigation({ waitUntil: 'networkidle2' })
      ]);
    } catch (err) {
      // Handle AJAX-loaded pages (no full reload): wait for new content to appear
      await page.waitForSelector('.item-class', { timeout: 5000 });
      console.log('Loaded next page via AJAX');
    }
  }

  // Final output: all scraped data
  console.log(`Total scraped items: ${allScrapedData.length}`);
  console.log(allScrapedData);

  await browser.close();
})();

Critical Notes for Success:

  • Use correct selectors: Open your target site in Chrome DevTools (F12) to find the exact CSS selector for your items and next page button.
  • Handle AJAX vs full reloads: Some sites load next pages via AJAX (no URL change), so waitForNavigation won’t work—use waitForSelector to wait for new content instead.
  • Add delays if needed: To avoid triggering anti-scraping measures, add await page.waitForTimeout(1000) after clicking the next button.
  • Debug with headless mode off: Setting headless: false lets you watch the browser in action, making it easy to spot why clicks or scraping aren’t working.

内容的提问来源于stack exchange,提问作者anthonyjdella

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:03:42