You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Puppeteer拦截window.location跳转并获取原始页面内容?

问题

使用Puppeteer抓取页面内容时,普通页面可以正常获取,但遇到通过window.location实现跳转的页面时,无法拦截跳转并获取原始HTML内容。例如访问https://example.com/thisredirects,页面返回包含跳转脚本的HTML,需要获取该HTML同时阻止跳转。

尝试用setRequestInterception拦截跳转,返回的response为null,且无法阻止这类跳转(该方法仅对HTTP状态码触发的跳转有效,对200响应后通过JS触发的window.location跳转无效),相关代码如下:

const puppeteer = require('puppeteer');

(async () => {
  const browser = await puppeteer.launch({ headless: false });
  const pageUrl = "https://example.com/thisredirects";

  const page = await browser.newPage();
  await page.setCacheEnabled(false);
  await page.setRequestInterception(true);

  const requests = [];
  page.on('request', async request => {
    let isNavRequest = request.isNavigationRequest() && request.frame() === page.mainFrame();
    if (!isNavRequest) {
      request.continue();
      return;
    }
    requests.push(request);
    if (requests.length == 1) {
      console.log("Load initial page: " + request.url());
      request.continue();
      return;
    }
    console.log("Block redirect to: " + request.url());
    request.abort();
  });

  let response;
  try {
    console.log(`Request: ${pageUrl}`);
    response = await page.goto(pageUrl, { waitUntil: 'domcontentloaded' });
    const content = await response.text();
    console.log(content);
    await page.close();
    await browser.close();
  }
  catch (err) {
    console.log(err);
  }
})()

另外尝试监听所有响应:

page.on('response', async response => {
  if (response.ok && response.url() === pageUrl) {
    console.log(await response.text());
  }
});

也无法获取原始HTML,抛出错误:无法加载此请求的内容。这可能是预检请求导致的。

请问有没有无需完全禁用JavaScript,就能拦截window.location跳转并获取原始HTML的方法?


解决方案

可以通过在页面加载前注入脚本重写window.location相关方法,拦截JS跳转,同时在跳转触发前获取原始页面内容,具体实现如下:

1. 核心思路

setRequestInterception只能拦截浏览器发起的导航请求,无法拦截页面内JS触发的window.location跳转。因此需要在页面脚本执行前,重写window.location的href setter以及assign、replace方法,阻止跳转行为,同时保留原始页面内容的可访问性。

2. 完整代码

const puppeteer = require('puppeteer');

(async () => {
  const browser = await puppeteer.launch({ headless: false });
  const pageUrl = "https://example.com/thisredirects";

  const page = await browser.newPage();
  await page.setCacheEnabled(false);

  // 在页面加载前注入拦截脚本
  await page.evaluateOnNewDocument(() => {
    // 保存原始location对象
    const originalLocation = window.location;
    
    // 重写href属性的setter,阻止直接赋值跳转
    Object.defineProperty(window, 'location', {
      get: () => originalLocation,
      set: (newUrl) => {
        console.log(`拦截window.location跳转: ${newUrl}`);
        // 不执行跳转逻辑
        return;
      },
      configurable: true
    });

    // 重写assign和replace方法,阻止跳转
    window.location.assign = (url) => {
      console.log(`拦截location.assign跳转: ${url}`);
    };
    window.location.replace = (url) => {
      console.log(`拦截location.replace跳转: ${url}`);
    };
  });

  try {
    console.log(`请求页面: ${pageUrl}`);
    // 使用domcontentloaded事件,确保在跳转脚本执行前获取内容
    await page.goto(pageUrl, { waitUntil: 'domcontentloaded' });
    
    // 获取原始页面的完整HTML
    const originalHtml = await page.content();
    console.log("原始页面内容:", originalHtml);

    await page.close();
    await browser.close();
  } catch (err) {
    console.error("错误:", err);
  }
})();

3. 关键说明

  • page.evaluateOnNewDocument:这个方法会在页面的任何脚本执行前注入代码,确保重写逻辑优先生效,不会被页面自身的脚本覆盖。
  • 选择waitUntil: 'domcontentloaded':这个事件在DOM加载完成后触发,此时原始HTML已经渲染完成,但页面内的跳转脚本可能还未执行(或刚触发就被拦截),可以保证成功获取原始内容。
  • 为什么之前监听response失败?当页面跳转触发后,浏览器会释放原始响应的资源,导致无法读取响应内容;阻止跳转后,原始响应的资源会被保留,就能正常获取。

内容的提问来源于stack exchange,提问作者Zak123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 15:27:48