You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Puppeteer递归获取多层嵌套iframe内容

使用Puppeteer获取多层嵌套iframe内容

问题背景

存在多层嵌套的iframe结构,示例如下:

top.html

<html>
  <title>top</title>
  <body>
    <p>top text</p>
    <iframe src="1.html"></iframe>
    <hr />
    <iframe src="2.html"></iframe>
  </body>
</html>

1.html

<html>
  <title>1</title>
  <body>
    <p>1 text</p>
    <iframe src="1-1.html"></iframe>
  </body>
</html>

1-1.html

<html>
  <title>1-1</title>
  <body>
    <p>1-1 text</p>
  </body>
</html>

2.html

<html>
  <title>2</title>
  <body>
    <p>2 text</p>
  </body>
</html>

目标是将所有嵌套iframe的内容展开合并为单一HTML字符串,可保留完整结构或仅提取核心内容(如所有<p>标签)。现有代码仅能处理单层iframe,递归实现因对evaluate的上下文差异理解不清而失败。

核心问题解析

Puppeteer的evaluate方法存在两个独立上下文:

  • Node.js上下文:运行Puppeteer API(如page.$$、contentFrame())
  • 浏览器上下文:evaluate的回调函数在浏览器环境中执行,可直接操作DOM

之前的单层处理逻辑在Node上下文操作iframe,递归时频繁跨上下文通信易出错。更简便的方式是将递归逻辑完全放在浏览器上下文执行,直接操作DOM和iframe的contentDocument。

解决方案:浏览器上下文递归替换iframe

以下代码可递归处理所有嵌套iframe,将内容展开后合并到主页面:

import { launch } from 'puppeteer';

(async () => {
  const browser = await launch({
    headless: 'new',
    args: [
      '--disable-web-security',
      '--disable-features=IsolateOrigins,site-per-process'
    ]
  });
  const page = await browser.newPage();
  // 加载目标页面,等待网络空闲确保iframe加载完成
  await page.goto('file:///C:/test/src/top.html', { waitUntil: 'networkidle0' });

  // 递归替换iframe的逻辑在浏览器上下文执行
  await page.evaluate(async () => {
    const replaceIframe = async (iframe) => {
      // 等待iframe加载完成
      await new Promise(resolve => {
        if (iframe.contentDocument?.readyState === 'complete') {
          resolve();
        } else {
          iframe.onload = resolve;
        }
      });

      const contentDoc = iframe.contentDocument;
      if (!contentDoc) return;

      // 先递归处理当前iframe内部的子iframe
      const childIframes = contentDoc.querySelectorAll('iframe');
      for (const childIframe of childIframes) {
        await replaceIframe(childIframe);
      }

      // 将iframe的完整内容插入到当前位置,移除原iframe
      const contentHtml = contentDoc.documentElement.outerHTML;
      iframe.insertAdjacentHTML('afterend', contentHtml);
      iframe.remove();
    };

    // 处理页面中所有顶级iframe
    const topIframes = document.querySelectorAll('iframe');
    for (const iframe of topIframes) {
      await replaceIframe(iframe);
    }
  });

  // 获取合并后的完整HTML
  const mergedHtml = await page.content();
  console.log(mergedHtml);

  await browser.close();
})();

简化需求:仅提取核心内容

若无需保留完整结构,仅需提取所有<p>标签内容,可修改递归逻辑:

import { launch } from 'puppeteer';

(async () => {
  const browser = await launch({
    headless: 'new',
    args: [
      '--disable-web-security',
      '--disable-features=IsolateOrigins,site-per-process'
    ]
  });
  const page = await browser.newPage();
  await page.goto('file:///C:/test/src/top.html', { waitUntil: 'networkidle0' });

  const extractedContent = await page.evaluate(async () => {
    const extractContent = async (doc) => {
      let content = '';
      // 提取当前文档的所有p标签
      doc.querySelectorAll('p').forEach(p => {
        content += `  ${p.outerHTML}\n\n`;
      });

      // 递归处理子iframe
      const iframes = doc.querySelectorAll('iframe');
      for (const iframe of iframes) {
        await new Promise(resolve => iframe.onload = resolve);
        if (iframe.contentDocument) {
          content += await extractContent(iframe.contentDocument);
        }
      }
      return content;
    };

    return await extractContent(document);
  });

  console.log(extractedContent);
  await browser.close();
})();

关键注意事项

  • 跨域场景必须保留启动参数中的--disable-web-security等配置,否则无法访问iframe的contentDocument
  • 需确保iframe加载完成后再处理,通过iframe.onload或readyState判断
  • 优先在浏览器上下文执行递归逻辑,避免Node与浏览器的上下文切换开销

内容的提问来源于stack exchange,提问作者enoeht

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 12:53:10