You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Node.js进行网页爬虫时如何绕过防火墙及规避请求拦截?

Node.js爬虫应对请求拦截与防火墙绕过方案

一、解决请求频率过高导致的拦截问题

这类拦截大多是网站反爬机制检测到异常请求频率,核心思路是让请求行为贴近正常用户:

  • 控制请求间隔并随机化:避免固定间隔的机械请求,在基础间隔上加入随机波动,降低被识别概率。比如基础间隔设为2秒,额外随机加0-1秒延迟:

    const axios = require('axios');
    
    async function fetchWithRandomDelay(url, baseDelay = 2000) {
      await new Promise(resolve => setTimeout(resolve, baseDelay + Math.random() * 1000));
      const response = await axios.get(url, {
        headers: {
          'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36',
          'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8'
        }
      });
      return response.data;
    }
    
  • 并发请求控制:不要一次性发起大量请求,限制同时请求数量,可用p-limit库实现:

    const pLimit = require('p-limit');
    const limit = pLimit(3); // 最多同时处理3个请求
    
    const urls = ['https://example.com/page1', 'https://example.com/page2', ...];
    const tasks = urls.map(url => limit(() => fetchWithRandomDelay(url)));
    
    Promise.all(tasks).then(results => {
      // 批量处理抓取结果
    });
    
  • 指数退避重试:遇到429(请求过多)、503(服务不可用)状态码时,按指数延长间隔后重试,避免加重服务器负担:

    async function fetchWithRetry(url, retries = 3, delay = 1000) {
      try {
        return await axios.get(url, { headers: { 'User-Agent': '...' } });
      } catch (err) {
        if (retries > 0 && [429, 503].includes(err.response?.status)) {
          await new Promise(resolve => setTimeout(resolve, delay));
          return fetchWithRetry(url, retries - 1, delay * 2); // 间隔翻倍
        }
        throw err;
      }
    }
    
  • 完善请求头:不要使用Node.js默认请求头,模拟浏览器的User-Agent、Accept、Referer等字段,让请求更像正常用户访问。

二、绕过防火墙的常用方案

防火墙拦截通常检测请求特征、IP、协议等,需针对性规避:

  • 轮换代理IP:使用代理池轮换IP,避免单一IP被封禁。可使用付费代理服务(稳定性更高),请求时循环切换代理:

    const proxyPool = [
      { host: 'proxy1.example.com', port: 8080 },
      { host: 'proxy2.example.com', port: 8080 },
      // 更多代理
    ];
    let proxyIndex = 0;
    
    async function fetchWithProxy(url) {
      const proxy = proxyPool[proxyIndex];
      proxyIndex = (proxyIndex + 1) % proxyPool.length; // 循环切换代理
      try {
        return await axios.get(url, {
          proxy: proxy,
          headers: { 'User-Agent': '...' }
        });
      } catch (err) {
        // 当前代理失效,切换下一个重试
        return fetchWithProxy(url);
      }
    }
    
  • 使用无头浏览器模拟真实行为:用Puppeteer或Playwright模拟点击、滚动、页面加载等待等用户操作,完全复刻浏览器请求:

    const puppeteer = require('puppeteer');
    
    async function fetchWithBrowser(url) {
      const browser = await puppeteer.launch({
        headless: 'new',
        args: ['--no-sandbox', '--disable-setuid-sandbox'] // 解决部分环境启动问题
      });
      const page = await browser.newPage();
      await page.setUserAgent('Mozilla/5.0 (Windows NT 10.0; Win64; x64) Chrome/114.0.0.0');
      await page.goto(url, { waitUntil: 'networkidle2' }); // 等待网络空闲,模拟用户加载完成
      const pageContent = await page.content();
      await browser.close();
      return pageContent;
    }
    
  • 启用HTTP/2协议:部分防火墙对HTTP/1.1特征检测更严格,使用Node.js原生http2模块发起请求:

    const http2 = require('http2');
    
    function fetchWithHttp2(url) {
      return new Promise((resolve, reject) => {
        const client = http2.connect(new URL(url).origin);
        client.on('error', reject);
    
        const req = client.request({
          ':path': new URL(url).pathname,
          'User-Agent': 'Mozilla/5.0 ...'
        });
    
        let data = '';
        req.on('data', chunk => data += chunk);
        req.on('end', () => {
          client.close();
          resolve(data);
        });
        req.end();
      });
    }
    
  • 维持会话一致性:保留并复用Cookie,模拟正常用户的会话持续行为,避免被防火墙判定为异常请求。

重要提示

所有爬虫行为必须遵守目标网站的robots.txt协议和服务条款,过度抓取或绕过反爬机制可能触发法律风险。

内容的提问来源于stack exchange,提问作者Burak Turna

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 08:30:52