Node.js Puppeteer网页爬虫翻页问题:异步await使用及按钮点击困惑
Hey there! Let's work through your Puppeteer crawler issues step by step—first fixing that await is only valid in async functions error, then tackling the pagination click logic.
First: Fix the await Error
The root cause here is straightforward: any line using await must live inside an async function. Since Puppeteer’s core methods (like launch(), goto(), evaluate()) all return Promises, you need to wrap your entire crawler logic in an async context.
Here’s the basic structure you should use (this eliminates the error immediately):
const puppeteer = require('puppeteer'); // Wrap everything in an async IIFE (Immediately Invoked Function Expression) (async () => { // Your crawler code goes here—all await calls will work now! const browser = await puppeteer.launch(); const page = await browser.newPage(); await page.goto('your-target-url'); // ... rest of your logic ... await browser.close(); })();
This is the standard pattern for Puppeteer scripts because it lets you use await without having to handle raw Promise .then() chains.
Second: Implement Reliable Pagination
Now for the tricky part: detecting and clicking the "next page" button repeatedly until there are no more pages. Let’s break this into a robust loop with safeguards:
Key Steps for Pagination:
- Initialize an array to store all scraped data
- Loop indefinitely until no "next page" button exists
- Scrape the current page’s data and add it to your array
- Check if the next page button is present
- If it exists: click it, wait for the page to load (handle both full page reloads and AJAX loads)
- If it doesn’t exist: exit the loop
Full Working Example
Replace the selectors (.item-class, .next-page-button) with ones that match your target website:
const puppeteer = require('puppeteer'); (async () => { // Launch browser with headless: false for easy debugging const browser = await puppeteer.launch({ headless: false }); const page = await browser.newPage(); // Navigate to the target page, wait for network to settle await page.goto('your-target-url', { waitUntil: 'networkidle2' }); const allScrapedData = []; while (true) { // Step 1: Scrape current page data (run code in browser context) const currentPageData = await page.evaluate(() => { // Replace this with your actual data extraction logic const items = document.querySelectorAll('.item-class'); return Array.from(items).map(item => ({ title: item.querySelector('.title').textContent.trim(), price: item.querySelector('.price').textContent.trim() // Add more fields as needed })); }); // Merge current page data into the main array allScrapedData.push(...currentPageData); console.log(`Scraped ${currentPageData.length} items from current page`); // Step 2: Check for next page button const nextPageButton = await page.$('.next-page-button'); if (!nextPageButton) { console.log('No more pages to scrape!'); break; // Exit loop when no next button exists } // Step 3: Click next page and wait for content to load try { // Handle full page reloads: wait for navigation to complete await Promise.all([ nextPageButton.click(), page.waitForNavigation({ waitUntil: 'networkidle2' }) ]); } catch (err) { // Handle AJAX-loaded pages (no full reload): wait for new content to appear await page.waitForSelector('.item-class', { timeout: 5000 }); console.log('Loaded next page via AJAX'); } } // Final output: all scraped data console.log(`Total scraped items: ${allScrapedData.length}`); console.log(allScrapedData); await browser.close(); })();
Critical Notes for Success:
- Use correct selectors: Open your target site in Chrome DevTools (F12) to find the exact CSS selector for your items and next page button.
- Handle AJAX vs full reloads: Some sites load next pages via AJAX (no URL change), so
waitForNavigationwon’t work—usewaitForSelectorto wait for new content instead. - Add delays if needed: To avoid triggering anti-scraping measures, add
await page.waitForTimeout(1000)after clicking the next button. - Debug with headless mode off: Setting
headless: falselets you watch the browser in action, making it easy to spot why clicks or scraping aren’t working.
内容的提问来源于stack exchange,提问作者anthonyjdella

