You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Puppeteer爬取Coles半价页面:已解决加载问题,求助请求追踪方法

Tracking Network Requests for Product Data in Puppeteer (AngularJS Site)

Hey there! Nice job resolving the initial page blank issue by using real browser cookies—those dynamic AngularJS sites can be tricky without the right session data. Let's dive into how you can track the network requests that are loading those half-price product listings, since you can see the data in the console but can't pinpoint the source request.

Core Approach: Listen to Network Responses/Requests

AngularJS sites typically load dynamic content via XHR (XMLHttpRequest) or Fetch API calls. Puppeteer lets you hook into these network events to capture the exact requests that return your product data. Here are a few actionable methods:

1. Track Responses and Filter for JSON Data

Add a response listener before navigating to the page. This will let you sniff every network response and extract JSON data from requests that likely contain product info:

const puppeteer = require('puppeteer');

async function getProductNames(){
  const browser = await puppeteer.launch({ headless: false });
  const page = await browser.newPage();
  await page.setViewport({ width: 1000, height: 926 });

  // 👇 Add this listener BEFORE page.goto
  page.on('response', async (response) => {
    // Filter URLs that might contain product data (adjust keywords to match the site)
    const url = response.url();
    if (url.includes('products') || url.includes('specials') || url.includes('half-price')) {
      try {
        // Parse the response as JSON (skip if it's not JSON)
        const data = await response.json();
        console.log('\n--- Found Product Data Request ---');
        console.log('URL:', url);
        console.log('Sample Product:', data?.products?.[0] || data?.items?.[0] || 'Check data structure');
        // You can save this data to a variable/file instead of logging
      } catch (err) {
        // Non-JSON response, ignore
      }
    }
  });

  await page.goto("https://shop.coles.com.au/a/richmond-south/specials/search/half-price-specials");
  await page.waitForSelector('.product-name');

  // Rest of your existing code...
  console.log("Begin to evaluate JS")
  var productNames = await page.evaluate(() => {
    const productElements = document.querySelectorAll('.product-name');
    return Array.from(productElements).map(el => el.textContent.trim());
  })
  console.log("Scraped Product Names:", productNames)
  
  await browser.close()
}
getProductNames();

2. Intercept Requests to Target Specific Resource Types

If you want to narrow down to only AJAX/Fetch requests (instead of all responses), use request interception to focus on xhr or fetch resource types:

async function getProductNames(){
  const browser = await puppeteer.launch({ headless: false });
  const page = await browser.newPage();
  await page.setViewport({ width: 1000, height: 926 });

  // Enable request interception
  await page.setRequestInterception(true);

  page.on('request', (request) => {
    // Only log XHR/Fetch requests (the ones that load dynamic data)
    if (['xhr', 'fetch'].includes(request.resourceType())) {
      console.log('AJAX Request:', request.url());
      // You can also check request headers or post data here if needed
    }
    request.continue(); // Don't block the request, just track it
  });

  // Rest of your code...
  await page.goto("https://shop.coles.com.au/a/richmond-south/specials/search/half-price-specials");
  // ...
}

3. Wait for a Specific Request to Complete

If you know roughly what the product data URL looks like, you can create a promise that resolves when that request finishes, so you can capture the full data directly:

async function getProductNames(){
  const browser = await puppeteer.launch({ headless: false });
  const page = await browser.newPage();
  await page.setViewport({ width: 1000, height: 926 });

  // Create a promise that waits for the product data response
  const productDataPromise = new Promise((resolve) => {
    page.on('response', async (response) => {
      // Adjust the URL match to what you see in the network tab (once you find it)
      if (response.url().includes('half-price-specials') && response.ok()) {
        const data = await response.json();
        resolve(data);
      }
    });
  });

  await page.goto("https://shop.coles.com.au/a/richmond-south/specials/search/half-price-specials");
  const productData = await productDataPromise;

  console.log("Full Product Data from API:", productData);
  // Extract names directly from the API data instead of scraping the DOM!
  const productNames = productData.products.map(p => p.name);
  console.log("Product Names from API:", productNames);
  
  await browser.close()
}

Pro Tips

  • Once you run the code, check the console for the URLs of requests returning product data. You can then inspect these URLs in Chrome's DevTools (while Puppeteer's non-headless browser is open) to see the exact request structure, headers, and response format.
  • If the site uses pagination, you'll need to trigger page navigation (e.g., click "Next" button) and continue listening for new responses as each page loads.
  • Since you're already using real browser cookies, you shouldn't hit authentication blocks for these API requests—they'll inherit the same session as the page.

内容的提问来源于stack exchange,提问作者Erron

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 16:47:49