You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

老旧tr/br/iframe架构网站爬取:iframe加载等待方案求助

Fixing iFrame Scraping with Reliable Waiting Mechanisms

The core issue with your current code is that setTimeout uses an arbitrary delay—this doesn’t account for variable load times (sometimes the iframe takes longer than 2s to load, or your code tries to access content before it’s ready). Additionally, your loop runs all iterations simultaneously, leading to race conditions where multiple clicks trigger iframe loads at once. Let’s fix this with two robust solutions:


Solution 1: Promise-Based Waiting Inside page.evaluate

This approach keeps logic within the browser context but replaces setTimeout with a promise that waits until the iframe content is fully loaded and your target element is available. We also make the loop sequential to avoid overlapping iframe requests.

const newResult = await page.evaluate(async (resultLength) => {
  const elements = Array.from(document.getElementsByClassName('class'));
  const results = [];

  // Iterate sequentially to avoid race conditions
  for (let i = 0; i < resultLength; i++) {
    const element = elements[i];
    const companyArray = element.innerHTML.split('<br>');
    let companyStreet, companyPostalCode;

    // Extract main page data
    const memberNumber = element.getElementsByTagName('a')[0].getAttribute('href').match(/[0-9]{1,5}/)[0];
    const companyName = companyArray[0]
      .replace(/<a[^>]*><span[^>]*><\/span>/, '')
      .replace(/<\/a>/, '');
    const companyNumber = companyArray[0].match(/[0-9]{6,8}/) 
      ? companyArray[0].match(/[0-9]{6,8}/)[0] 
      : '';
    const companyTown = companyArray[1].replace('"', '');
    const companyRegion = companyArray[2].replace(/<span[^>]*>Some text:<\/span>/, '');
    const telNumber = element.innerHTML
      .substring(element.innerHTML.lastIndexOf('</span>'))
      .replace('</span>', '')
      .replace('<br>', '');

    // Trigger iframe load
    element.getElementsByTagName('a')[0].click();

    // Wait for iframe content to be ready (custom promise)
    const waitForIFrameContent = () => {
      return new Promise((resolve) => {
        const iframe = document.getElementById('some-id');

        function checkForTargetElement() {
          const contentDoc = iframe.contentWindow.document;
          const targetElement = contentDoc.getElementById('lblAdresse');
          // Resolve only when element exists and has content
          if (targetElement && targetElement.innerHTML.trim() !== '') {
            resolve(targetElement.innerHTML);
          } else {
            setTimeout(checkForTargetElement, 100); // Retry every 100ms
          }
        }

        // If iframe is already loaded, check immediately; else wait for load event
        if (iframe.contentDocument.readyState === 'complete') {
          checkForTargetElement();
        } else {
          iframe.addEventListener('load', checkForTargetElement);
        }
      });
    };

    // Wait for iframe data before proceeding
    const iFrameContentHTML = await waitForIFrameContent();
    const iFrameContent = iFrameContentHTML.split('<br>');
    companyStreet = iFrameContent[0].replace('"', '');
    companyPostalCode = iFrameContent[2].replace('"', '');

    console.log(companyStreet, companyPostalCode);
    results.push({ memberNumber, companyName, companyNumber, companyTown, companyRegion, telNumber, companyStreet, companyPostalCode });
  }

  return results;
}, pageSearchResults.length);

Key improvements:

  • Sequential loop: Each iteration waits for the iframe to load before moving to the next entry.
  • Dynamic waiting: The promise checks for the presence of your target element (lblAdresse) with actual content, instead of relying on a fixed delay.
  • Result collection: Returns structured results instead of relying on side effects.

Solution 2: Use Puppeteer’s Frame API (More Robust)

For better control and reliability, leverage Puppeteer’s built-in frame handling outside of page.evaluate. This avoids browser context limitations and uses Puppeteer’s optimized waiting methods.

const results = [];
const elements = await page.$$('.class'); // Get all elements with the target class

for (let i = 0; i < elements.length; i++) {
  const element = elements[i];

  // Extract data from the main page
  const elementHTML = await element.evaluate(el => el.innerHTML);
  const companyArray = elementHTML.split('<br>');

  const memberNumber = await element.$eval('a', a => a.getAttribute('href').match(/[0-9]{1,5}/)[0]);
  const companyName = companyArray[0]
    .replace(/<a[^>]*><span[^>]*><\/span>/, '')
    .replace(/<\/a>/, '');
  const companyNumber = companyArray[0].match(/[0-9]{6,8}/) 
    ? companyArray[0].match(/[0-9]{6,8}/)[0] 
    : '';
  const companyTown = companyArray[1].replace('"', '');
  const companyRegion = companyArray[2].replace(/<span[^>]*>Some text:<\/span>/, '');
  const telNumber = elementHTML
    .substring(elementHTML.lastIndexOf('</span>'))
    .replace('</span>', '')
    .replace('<br>', '');

  // Trigger iframe load
  await element.$eval('a', a => a.click());

  // Wait for the iframe to be available and loaded
  const iframeElement = await page.waitForSelector('#some-id');
  const iframe = await iframeElement.contentFrame();

  // Wait for the target element inside the iframe
  await iframe.waitForSelector('#lblAdresse');
  const iFrameContentHTML = await iframe.$eval('#lblAdresse', el => el.innerHTML);
  const iFrameContent = iFrameContentHTML.split('<br>');

  const companyStreet = iFrameContent[0].replace('"', '');
  const companyPostalCode = iFrameContent[2].replace('"', '');

  console.log(companyStreet, companyPostalCode);
  results.push({ memberNumber, companyName, companyNumber, companyTown, companyRegion, telNumber, companyStreet, companyPostalCode });
}

Key advantages:

  • Puppeteer’s native waits: waitForSelector and contentFrame handle edge cases like slow network loads or dynamic content injection.
  • Better debugging: Easier to add breakpoints and inspect frame state.
  • Avoids cross-origin issues: Puppeteer handles same-origin frame access seamlessly (if allowed by the site’s CORS policy).

Additional Tips

  • Error handling: Wrap each iteration in a try/catch block to skip failed entries instead of crashing the entire scrape.
  • Iframe URL filtering: If you know the iframe’s expected URL, use page.waitForFrame(frame => frame.url().includes('your-iframe-url')) to target the correct frame.
  • Throttling: Add a small delay between iterations if the site has anti-scraping measures.

内容的提问来源于stack exchange,提问作者Jean-Baptiste

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:45:43