You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

请求目标院校页面始终无法获取200状态码,求非Selenium的高效爬取方案

请求目标院校页面始终无法获取200状态码,求非Selenium的高效爬取方案

兄弟,我之前爬类似的教育类门户也碰到过这种情况——明明看了robots.txt允许,headers也改了,就是拿不到200。大概率是网站用了一些隐蔽的反爬手段,不是单纯看表面的headers。给你几个非Selenium的高效方案,亲测有用:

  • 优化请求头的完整性,不要只改User-Agent
    很多网站会校验请求头的完整度,不是随便写个Chrome的UA就行。你要把真实浏览器发送的所有关键字段都带上,比如Accept、Accept-Language、Referer、Accept-Encoding这些。举个例子,完整的请求头可以这么写:

    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
        "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8",
        "Accept-Language": "en-US,en;q=0.5",
        "Referer": "https://www.mastersportal.com/",
        "Accept-Encoding": "gzip, deflate, br"
    }
    
  • 用支持模拟TLS指纹的库绕过检测
    现在很多网站会通过TLS握手的指纹识别爬虫,普通的requests库的TLS指纹和真实浏览器差异很大,很容易被拦截。推荐用curl_cffi这个库,它可以直接模拟Chrome/Firefox等浏览器的TLS指纹,代码示例:

    from curl_cffi import requests
    
    url = "https://www.mastersportal.com/universities/80/tilburg-university.html"
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
        "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8",
        "Accept-Language": "en-US,en;q=0.5",
        "Referer": "https://www.mastersportal.com/"
    }
    
    # 模拟Chrome 118的指纹和行为
    response = requests.get(url, headers=headers, impersonate="chrome118")
    print(response.status_code)
    

    这个方法我用过很多次,对付这种隐蔽的反爬特别有效,而且速度比Selenium快太多。

  • 尝试使用HTTP2协议请求
    不少现代网站已经强制或优先支持HTTP2,而requests默认用的是HTTP1.1,这也可能导致被拦截。可以用httpx库开启HTTP2支持试试:

    import httpx
    
    url = "https://www.mastersportal.com/universities/80/tilburg-university.html"
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
        "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8",
        "Accept-Language": "en-US,en;q=0.5",
        "Referer": "https://www.mastersportal.com/"
    }
    
    # 开启HTTP2支持
    client = httpx.Client(http2=True)
    response = client.get(url, headers=headers)
    print(response.status_code)
    
  • 控制请求频率,避免触发速率限制
    就算前面的都做好了,短时间内连续请求也可能被网站临时封禁。可以加个随机延迟,比如每次请求后暂停1-3秒:

    import random
    import time
    
    # 请求后延迟
    time.sleep(random.uniform(1, 3))
    

如果上面的方法都试过还是不行,你可以试试先请求网站的首页,获取会话Cookie后再请求目标页面——有些网站会给首次访问的用户设置会话标识,没有这个标识的请求会被拦截。

备注:内容来源于stack exchange,提问作者hassan abbas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 16:39:31