You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫遭拦截但浏览器可正常访问的问题咨询与解决需求

问题背景

尝试用Python爬取https://www.[cencored].com/,最初用简单的requests.get()能正常运行,次日失败(期间Windows进行了更新,不确定是否相关)。添加请求头后代码如下:

import requests

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'
}
print(requests.get(url, headers=headers).text)

出现错误:

requests.exceptions.ConnectionError: HTTPSConnectionPool(host='[cencored].com', port=443): Max retries exceeded with url: / (Caused by NewConnectionError('<urllib3.connection.HTTPSConnection object at 0x0000025316C12430>: Failed to establish a new connection: [WinError 10060] A connection attempt failed because the connected party did not properly respond after a period of time, or established connection failed because connected host has failed to respond'))

改用selenium后仍无法访问,返回:

502 Bad Gateway

ProtocolException('Server connection to ('[cencored].com', 443) failed: Error connecting to "xxxx.com": [WinError 10060] A connection attempt failed because the connected party did not properly respond after a period of time, or established connection failed because connected host has failed to respond')

确认IP未被拉黑(Google Chrome可正常访问该网站),有两个疑问:

问题
  1. Google Chrome为何不会被拦截?
  2. 我该如何绕过反爬保护?
解答

1. Google Chrome不被拦截的原因

  • 请求头完整性:Chrome发送的请求头远不止User-Agent,还包含Accept、Accept-Language、Accept-Encoding、Referer、Cookie、Connection等多个字段,这些字段组合起来更接近真实用户的请求特征,反爬系统会判定为正常访问。
  • 会话与Cookie:Chrome会保存网站的Cookie(包括会话Cookie、登录状态Cookie等),很多反爬系统会验证Cookie有效性,而你的requests/selenium请求可能未携带这些有效Cookie。
  • 请求频率与行为模式:Chrome的访问频率是人为操作的,间隔时间随机且符合正常用户习惯;而爬虫请求可能短时间内发送大量请求,或请求模式过于机械,触发IP级拦截,但Chrome因之前的正常会话被允许访问。
  • TLS/SSL指纹差异:Chrome的TLS握手参数、SSL加密套件等,和requests/selenium使用的底层库(如urllib3、Chromedriver)有差异,反爬系统可通过这些指纹识别非浏览器请求并拦截。

2. 绕过反爬保护的可行方法

  • 完善请求头:复制Chrome请求的完整请求头(所有字段)到requests中,不要只加User-Agent。可在Chrome开发者工具的Network面板中,找到对应请求的Request Headers全部复制。
  • 复用浏览器Cookie:从Chrome中导出当前网站的Cookie,添加到requests的headers或cookies参数中,模拟已建立的会话。
  • 控制请求频率:添加随机延迟(比如time.sleep(random.uniform(1,3))),避免短时间内发送大量请求,模拟人类操作间隔。
  • 使用代理IP:如果确实是IP被针对,更换代理IP(优先选择付费代理池,稳定性更高),每次请求轮换不同IP。
  • 优化Selenium配置:
    • 禁用Chromedriver的特征标识,比如添加--disable-blink-features=AutomationControlled参数,避免被检测出是自动化工具。
    • 模拟人类行为,比如随机滚动页面、点击元素、等待页面加载完成后再获取内容,而非直接抓取页面源码。
    • 使用无头模式时,添加更多模拟真实浏览器的参数,比如设置窗口大小、启用图片加载等。
  • 使用浏览器指纹模拟工具:用fake_useragent库随机生成真实的User-Agent,或配合requests-futures等工具,让请求指纹更接近真实浏览器。

内容的提问来源于stack exchange,提问作者Electron X

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 15:43:16