如何使用Puppeteer递归获取多层嵌套iframe内容
使用Puppeteer获取多层嵌套iframe内容
问题背景
存在多层嵌套的iframe结构,示例如下:
top.html
<html> <title>top</title> <body> <p>top text</p> <iframe src="1.html"></iframe> <hr /> <iframe src="2.html"></iframe> </body> </html>
1.html
<html> <title>1</title> <body> <p>1 text</p> <iframe src="1-1.html"></iframe> </body> </html>
1-1.html
<html> <title>1-1</title> <body> <p>1-1 text</p> </body> </html>
2.html
<html> <title>2</title> <body> <p>2 text</p> </body> </html>
目标是将所有嵌套iframe的内容展开合并为单一HTML字符串,可保留完整结构或仅提取核心内容(如所有<p>标签)。现有代码仅能处理单层iframe,递归实现因对evaluate的上下文差异理解不清而失败。
核心问题解析
Puppeteer的evaluate方法存在两个独立上下文:
- Node.js上下文:运行Puppeteer API(如
page.$$、contentFrame()) - 浏览器上下文:
evaluate的回调函数在浏览器环境中执行,可直接操作DOM
之前的单层处理逻辑在Node上下文操作iframe,递归时频繁跨上下文通信易出错。更简便的方式是将递归逻辑完全放在浏览器上下文执行,直接操作DOM和iframe的contentDocument。
解决方案:浏览器上下文递归替换iframe
以下代码可递归处理所有嵌套iframe,将内容展开后合并到主页面:
import { launch } from 'puppeteer'; (async () => { const browser = await launch({ headless: 'new', args: [ '--disable-web-security', '--disable-features=IsolateOrigins,site-per-process' ] }); const page = await browser.newPage(); // 加载目标页面,等待网络空闲确保iframe加载完成 await page.goto('file:///C:/test/src/top.html', { waitUntil: 'networkidle0' }); // 递归替换iframe的逻辑在浏览器上下文执行 await page.evaluate(async () => { const replaceIframe = async (iframe) => { // 等待iframe加载完成 await new Promise(resolve => { if (iframe.contentDocument?.readyState === 'complete') { resolve(); } else { iframe.onload = resolve; } }); const contentDoc = iframe.contentDocument; if (!contentDoc) return; // 先递归处理当前iframe内部的子iframe const childIframes = contentDoc.querySelectorAll('iframe'); for (const childIframe of childIframes) { await replaceIframe(childIframe); } // 将iframe的完整内容插入到当前位置,移除原iframe const contentHtml = contentDoc.documentElement.outerHTML; iframe.insertAdjacentHTML('afterend', contentHtml); iframe.remove(); }; // 处理页面中所有顶级iframe const topIframes = document.querySelectorAll('iframe'); for (const iframe of topIframes) { await replaceIframe(iframe); } }); // 获取合并后的完整HTML const mergedHtml = await page.content(); console.log(mergedHtml); await browser.close(); })();
简化需求:仅提取核心内容
若无需保留完整结构,仅需提取所有<p>标签内容,可修改递归逻辑:
import { launch } from 'puppeteer'; (async () => { const browser = await launch({ headless: 'new', args: [ '--disable-web-security', '--disable-features=IsolateOrigins,site-per-process' ] }); const page = await browser.newPage(); await page.goto('file:///C:/test/src/top.html', { waitUntil: 'networkidle0' }); const extractedContent = await page.evaluate(async () => { const extractContent = async (doc) => { let content = ''; // 提取当前文档的所有p标签 doc.querySelectorAll('p').forEach(p => { content += ` ${p.outerHTML}\n\n`; }); // 递归处理子iframe const iframes = doc.querySelectorAll('iframe'); for (const iframe of iframes) { await new Promise(resolve => iframe.onload = resolve); if (iframe.contentDocument) { content += await extractContent(iframe.contentDocument); } } return content; }; return await extractContent(document); }); console.log(extractedContent); await browser.close(); })();
关键注意事项
- 跨域场景必须保留启动参数中的
--disable-web-security等配置,否则无法访问iframe的contentDocument - 需确保iframe加载完成后再处理,通过
iframe.onload或readyState判断 - 优先在浏览器上下文执行递归逻辑,避免Node与浏览器的上下文切换开销
内容的提问来源于stack exchange,提问作者enoeht
相关产品推荐
相关产品推荐

