Puppeteer中response.status()无法识别300系列状态码求助
解决Puppeteer无法检测301重定向状态码的问题
你的问题出在page.goto()的默认行为上:它会自动跟随所有重定向,返回的是最终跳转完成后页面的响应,所以你拿到的是重定向目标页面的200状态码,而非原始链接的301。
下面提供两种可行的解决方案:
方案一:启用请求拦截追踪重定向链
通过Puppeteer的请求拦截功能,捕获每个请求的响应状态,包括重定向过程中的301:
async function categorizeLinks(req, res, url) { const browser = await puppeteer.launch({ headless: 'false', args: ['--no-sandbox'] }) const page = await browser.newPage(); await page.setViewport({ width: 1366, height: 1068 }); await page.goto(url); // Extract all links from the page const links = await page.$$eval('a', (anchors) => { return anchors.map((anchor) => anchor.href); }); console.log(links); const categorizedLinks = { '200': [], '300': [], '400': [], '500': [], }; // 开启请求拦截 await page.setRequestInterception(true); page.on('request', (request) => request.continue()); for (const link of links) { try { let redirectStatus = null; // 监听响应链,优先记录重定向状态 const responsePromise = new Promise(resolve => { const handleResponse = (response) => { const status = response.status(); if (status >= 300 && status < 400) { // 记录第一个重定向状态 redirectStatus = status; // 继续监听下一个响应(重定向目标) page.once('response', handleResponse); } else { // 返回最终状态,若有重定向则用重定向状态 resolve(redirectStatus || status); } }; page.once('response', handleResponse); }); await page.goto(link, { timeout: 10000 }); const statusCode = await responsePromise; // 分类链接 if (statusCode >= 200 && statusCode <= 299) { categorizedLinks['200'].push(link); } else if (statusCode >= 300 && statusCode <= 399) { categorizedLinks['300'].push(link); } else if (statusCode >= 400 && statusCode <= 499) { categorizedLinks['400'].push(link); } else if (statusCode >= 500) { categorizedLinks['500'].push(link); } } catch (error) { console.error(`Failed to fetch link: ${link}`, error); } } await browser.close(); console.log("categorizedLinks", categorizedLinks); }
方案二:用原生HTTP请求替代页面跳转(更高效)
如果不需要加载页面内容,只是要获取状态码,直接用Node.js的fetch发送请求,禁止自动跟随重定向,直接拿到原始状态码:
async function categorizeLinks(req, res, url) { const browser = await puppeteer.launch({ headless: 'false', args: ['--no-sandbox'] }) const page = await browser.newPage(); await page.setViewport({ width: 1366, height: 1068 }); await page.goto(url); // Extract all links from the page const links = await page.$$eval('a', (anchors) => { return anchors.map((anchor) => anchor.href); }); console.log(links); const categorizedLinks = { '200': [], '300': [], '400': [], '500': [], }; // 用fetch替代page.goto,禁止自动跟随重定向 for (const link of links) { try { const response = await fetch(link, { method: 'GET', redirect: 'manual', // 关键:不跟随重定向 signal: AbortSignal.timeout(10000) }); const statusCode = response.status; // 分类链接 if (statusCode >= 200 && statusCode <= 299) { categorizedLinks['200'].push(link); } else if (statusCode >= 300 && statusCode <= 399) { categorizedLinks['300'].push(link); } else if (statusCode >= 400 && statusCode <= 499) { categorizedLinks['400'].push(link); } else if (statusCode >= 500) { categorizedLinks['500'].push(link); } } catch (error) { console.error(`Failed to fetch link: ${link}`, error); } } await browser.close(); console.log("categorizedLinks", categorizedLinks); }
注意事项
- 方案二不需要频繁跳转页面,执行效率更高,适合只需要状态码的场景。
- 如果遇到跨域限制导致
fetch失败,建议切换回方案一。 - 页面中提取的
href已经是绝对路径,无需额外处理相对路径问题。
内容的提问来源于stack exchange,提问作者Jason Salgado
相关产品推荐
相关产品推荐

