请求返回502状态码,网页爬取无法获取原始HTML求助
解决爬取https://und.com时502错误的方案
一、核心解决502状态码问题
502 Bad Gateway大多是目标网站反爬机制识别出你的请求为非浏览器请求,直接拒绝访问。解决方法是给Axios请求添加浏览器标准请求头,模拟真实用户访问:
修改Axios请求配置,补充必要的请求头:
axios({ url: website, headers: { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Connection': 'keep-alive', 'Upgrade-Insecure-Requests': '1' } })
二、修复异步错误处理漏洞
你原代码里的try/catch无法捕获Axios Promise的rejection(因为axios().then()是异步操作),必须改用.catch()或async/await处理:
方式1:用.catch()捕获异步错误
axios({ /* 上述请求配置 */ }) .then((res) => { // 处理响应逻辑 }) .catch(error => { console.error('请求失败:', error.response?.status, error.message); });
方式2:改用async/await配合try/catch
async function fetchSiteData() { try { const res = await axios({ /* 上述请求配置 */ }); const $ = cheerio.load(res.data); let content = []; $(".sdc-site-tile__headline").each(function () { const title = $(this).text().trim(); const url = $(this).find("a").attr("href"); // 补全相对URL为绝对地址 const fullUrl = url.startsWith('http') ? url : `${website}${url}`; content.push({ title, url: fullUrl }); }); // 统一注册路由 app.get("/", (req, res) => { res.json(content); }); } catch (error) { console.error('爬取失败:', error.response?.status, error.message); // 路由返回错误信息 app.get("/", (req, res) => { res.status(500).json({ error: '爬取数据失败', details: error.message }); }); } } fetchSiteData();
三、修复路由注册的错误
你原代码把app.get("/", ...)放在.each()循环里,会导致重复注册多个相同路由,最终只有最后一个生效。必须把路由注册逻辑放在循环外,或者直接在路由触发时执行爬取(适合需要实时数据的场景):
app.get("/", async (req, res) => { try { const res = await axios({ /* 请求配置 */ }); const $ = cheerio.load(res.data); let content = []; $(".sdc-site-tile__headline").each(function () { const title = $(this).text().trim(); const url = $(this).find("a").attr("href"); const fullUrl = url.startsWith('http') ? url : `${website}${url}`; content.push({ title, url: fullUrl }); }); res.json(content); } catch (error) { res.status(500).json({ error: '获取数据失败', status: error.response?.status }); } });
四、其他备选方案
- 如果添加请求头后仍返回502,可能需要使用代理IP避免IP被目标网站封禁。
- 先查看目标网站的
robots.txt,确认是否允许爬取相关内容。
内容的提问来源于stack exchange,提问作者rpuh
相关产品推荐
相关产品推荐

