如何用Puppeteer拦截window.location跳转并获取原始页面内容?
问题
使用Puppeteer抓取页面内容时,普通页面可以正常获取,但遇到通过window.location实现跳转的页面时,无法拦截跳转并获取原始HTML内容。例如访问https://example.com/thisredirects,页面返回包含跳转脚本的HTML,需要获取该HTML同时阻止跳转。
尝试用setRequestInterception拦截跳转,返回的response为null,且无法阻止这类跳转(该方法仅对HTTP状态码触发的跳转有效,对200响应后通过JS触发的window.location跳转无效),相关代码如下:
const puppeteer = require('puppeteer'); (async () => { const browser = await puppeteer.launch({ headless: false }); const pageUrl = "https://example.com/thisredirects"; const page = await browser.newPage(); await page.setCacheEnabled(false); await page.setRequestInterception(true); const requests = []; page.on('request', async request => { let isNavRequest = request.isNavigationRequest() && request.frame() === page.mainFrame(); if (!isNavRequest) { request.continue(); return; } requests.push(request); if (requests.length == 1) { console.log("Load initial page: " + request.url()); request.continue(); return; } console.log("Block redirect to: " + request.url()); request.abort(); }); let response; try { console.log(`Request: ${pageUrl}`); response = await page.goto(pageUrl, { waitUntil: 'domcontentloaded' }); const content = await response.text(); console.log(content); await page.close(); await browser.close(); } catch (err) { console.log(err); } })()
另外尝试监听所有响应:
page.on('response', async response => { if (response.ok && response.url() === pageUrl) { console.log(await response.text()); } });
也无法获取原始HTML,抛出错误:无法加载此请求的内容。这可能是预检请求导致的。
请问有没有无需完全禁用JavaScript,就能拦截window.location跳转并获取原始HTML的方法?
解决方案
可以通过在页面加载前注入脚本重写window.location相关方法,拦截JS跳转,同时在跳转触发前获取原始页面内容,具体实现如下:
1. 核心思路
setRequestInterception只能拦截浏览器发起的导航请求,无法拦截页面内JS触发的window.location跳转。因此需要在页面脚本执行前,重写window.location的href setter以及assign、replace方法,阻止跳转行为,同时保留原始页面内容的可访问性。
2. 完整代码
const puppeteer = require('puppeteer'); (async () => { const browser = await puppeteer.launch({ headless: false }); const pageUrl = "https://example.com/thisredirects"; const page = await browser.newPage(); await page.setCacheEnabled(false); // 在页面加载前注入拦截脚本 await page.evaluateOnNewDocument(() => { // 保存原始location对象 const originalLocation = window.location; // 重写href属性的setter,阻止直接赋值跳转 Object.defineProperty(window, 'location', { get: () => originalLocation, set: (newUrl) => { console.log(`拦截window.location跳转: ${newUrl}`); // 不执行跳转逻辑 return; }, configurable: true }); // 重写assign和replace方法,阻止跳转 window.location.assign = (url) => { console.log(`拦截location.assign跳转: ${url}`); }; window.location.replace = (url) => { console.log(`拦截location.replace跳转: ${url}`); }; }); try { console.log(`请求页面: ${pageUrl}`); // 使用domcontentloaded事件,确保在跳转脚本执行前获取内容 await page.goto(pageUrl, { waitUntil: 'domcontentloaded' }); // 获取原始页面的完整HTML const originalHtml = await page.content(); console.log("原始页面内容:", originalHtml); await page.close(); await browser.close(); } catch (err) { console.error("错误:", err); } })();
3. 关键说明
page.evaluateOnNewDocument:这个方法会在页面的任何脚本执行前注入代码,确保重写逻辑优先生效,不会被页面自身的脚本覆盖。- 选择
waitUntil: 'domcontentloaded':这个事件在DOM加载完成后触发,此时原始HTML已经渲染完成,但页面内的跳转脚本可能还未执行(或刚触发就被拦截),可以保证成功获取原始内容。 - 为什么之前监听
response失败?当页面跳转触发后,浏览器会释放原始响应的资源,导致无法读取响应内容;阻止跳转后,原始响应的资源会被保留,就能正常获取。
内容的提问来源于stack exchange,提问作者Zak123
相关产品推荐
相关产品推荐

