老旧tr/br/iframe架构网站爬取:iframe加载等待方案求助
The core issue with your current code is that setTimeout uses an arbitrary delay—this doesn’t account for variable load times (sometimes the iframe takes longer than 2s to load, or your code tries to access content before it’s ready). Additionally, your loop runs all iterations simultaneously, leading to race conditions where multiple clicks trigger iframe loads at once. Let’s fix this with two robust solutions:
Solution 1: Promise-Based Waiting Inside page.evaluate
This approach keeps logic within the browser context but replaces setTimeout with a promise that waits until the iframe content is fully loaded and your target element is available. We also make the loop sequential to avoid overlapping iframe requests.
const newResult = await page.evaluate(async (resultLength) => { const elements = Array.from(document.getElementsByClassName('class')); const results = []; // Iterate sequentially to avoid race conditions for (let i = 0; i < resultLength; i++) { const element = elements[i]; const companyArray = element.innerHTML.split('<br>'); let companyStreet, companyPostalCode; // Extract main page data const memberNumber = element.getElementsByTagName('a')[0].getAttribute('href').match(/[0-9]{1,5}/)[0]; const companyName = companyArray[0] .replace(/<a[^>]*><span[^>]*><\/span>/, '') .replace(/<\/a>/, ''); const companyNumber = companyArray[0].match(/[0-9]{6,8}/) ? companyArray[0].match(/[0-9]{6,8}/)[0] : ''; const companyTown = companyArray[1].replace('"', ''); const companyRegion = companyArray[2].replace(/<span[^>]*>Some text:<\/span>/, ''); const telNumber = element.innerHTML .substring(element.innerHTML.lastIndexOf('</span>')) .replace('</span>', '') .replace('<br>', ''); // Trigger iframe load element.getElementsByTagName('a')[0].click(); // Wait for iframe content to be ready (custom promise) const waitForIFrameContent = () => { return new Promise((resolve) => { const iframe = document.getElementById('some-id'); function checkForTargetElement() { const contentDoc = iframe.contentWindow.document; const targetElement = contentDoc.getElementById('lblAdresse'); // Resolve only when element exists and has content if (targetElement && targetElement.innerHTML.trim() !== '') { resolve(targetElement.innerHTML); } else { setTimeout(checkForTargetElement, 100); // Retry every 100ms } } // If iframe is already loaded, check immediately; else wait for load event if (iframe.contentDocument.readyState === 'complete') { checkForTargetElement(); } else { iframe.addEventListener('load', checkForTargetElement); } }); }; // Wait for iframe data before proceeding const iFrameContentHTML = await waitForIFrameContent(); const iFrameContent = iFrameContentHTML.split('<br>'); companyStreet = iFrameContent[0].replace('"', ''); companyPostalCode = iFrameContent[2].replace('"', ''); console.log(companyStreet, companyPostalCode); results.push({ memberNumber, companyName, companyNumber, companyTown, companyRegion, telNumber, companyStreet, companyPostalCode }); } return results; }, pageSearchResults.length);
Key improvements:
- Sequential loop: Each iteration waits for the iframe to load before moving to the next entry.
- Dynamic waiting: The promise checks for the presence of your target element (
lblAdresse) with actual content, instead of relying on a fixed delay. - Result collection: Returns structured results instead of relying on side effects.
Solution 2: Use Puppeteer’s Frame API (More Robust)
For better control and reliability, leverage Puppeteer’s built-in frame handling outside of page.evaluate. This avoids browser context limitations and uses Puppeteer’s optimized waiting methods.
const results = []; const elements = await page.$$('.class'); // Get all elements with the target class for (let i = 0; i < elements.length; i++) { const element = elements[i]; // Extract data from the main page const elementHTML = await element.evaluate(el => el.innerHTML); const companyArray = elementHTML.split('<br>'); const memberNumber = await element.$eval('a', a => a.getAttribute('href').match(/[0-9]{1,5}/)[0]); const companyName = companyArray[0] .replace(/<a[^>]*><span[^>]*><\/span>/, '') .replace(/<\/a>/, ''); const companyNumber = companyArray[0].match(/[0-9]{6,8}/) ? companyArray[0].match(/[0-9]{6,8}/)[0] : ''; const companyTown = companyArray[1].replace('"', ''); const companyRegion = companyArray[2].replace(/<span[^>]*>Some text:<\/span>/, ''); const telNumber = elementHTML .substring(elementHTML.lastIndexOf('</span>')) .replace('</span>', '') .replace('<br>', ''); // Trigger iframe load await element.$eval('a', a => a.click()); // Wait for the iframe to be available and loaded const iframeElement = await page.waitForSelector('#some-id'); const iframe = await iframeElement.contentFrame(); // Wait for the target element inside the iframe await iframe.waitForSelector('#lblAdresse'); const iFrameContentHTML = await iframe.$eval('#lblAdresse', el => el.innerHTML); const iFrameContent = iFrameContentHTML.split('<br>'); const companyStreet = iFrameContent[0].replace('"', ''); const companyPostalCode = iFrameContent[2].replace('"', ''); console.log(companyStreet, companyPostalCode); results.push({ memberNumber, companyName, companyNumber, companyTown, companyRegion, telNumber, companyStreet, companyPostalCode }); }
Key advantages:
- Puppeteer’s native waits:
waitForSelectorandcontentFramehandle edge cases like slow network loads or dynamic content injection. - Better debugging: Easier to add breakpoints and inspect frame state.
- Avoids cross-origin issues: Puppeteer handles same-origin frame access seamlessly (if allowed by the site’s CORS policy).
Additional Tips
- Error handling: Wrap each iteration in a
try/catchblock to skip failed entries instead of crashing the entire scrape. - Iframe URL filtering: If you know the iframe’s expected URL, use
page.waitForFrame(frame => frame.url().includes('your-iframe-url'))to target the correct frame. - Throttling: Add a small delay between iterations if the site has anti-scraping measures.
内容的提问来源于stack exchange,提问作者Jean-Baptiste

