You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium为Discord机器人爬取WarThunder玩家信息页面时遭遇Cloudflare无限加载的技术求助

解决Selenium访问WarThunder玩家页面时Cloudflare无限加载的问题

我之前也碰到过类似的Cloudflare反爬卡点,结合你的情况来看,核心原因就是Selenium自带的自动化特征被Cloudflare精准识别了。下面给你几个亲测有效的解决方案,按落地优先级排序:

1. 用ChromeOptions彻底抹除WebDriver痕迹

Cloudflare现在会直接通过navigator.webdriver这个浏览器属性识别Selenium,所以第一步就要把这个标志彻底隐藏,同时搭配一堆模拟真实浏览器的参数:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options

options = Options()
# 核心:禁用自动化检测开关
options.add_argument("--disable-blink-features=AutomationControlled")
# 替换成你自己常用浏览器的UA(可以在浏览器控制台输入navigator.userAgent获取)
options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")
# 禁用扩展、沙箱,避免额外特征暴露
options.add_argument("--disable-extensions")
options.add_argument("--no-sandbox")
options.add_argument("--disable-dev-shm-usage")
# 最大化窗口,模拟真实用户打开浏览器的默认状态
options.add_argument("--start-maximized")

# 初始化驱动时传入配置
driver = webdriver.Chrome(options=options)

2. 使用专门的反检测驱动:undetected-chromedriver

这个库是针对Chrome的修改版驱动,会自动处理Cloudflare的大部分检测逻辑,比手动调参数省心太多:

先安装库:

pip install undetected-chromedriver

然后直接使用:

import undetected_chromedriver as uc

# 初始化驱动,默认会自动隐藏所有自动化特征
driver = uc.Chrome()
driver.get("https://warthunder.com/en/community/userinfo/?nick=Hunter_i86")
# 加个隐式等待,确保页面完全加载
driver.implicitly_wait(10)

我用这个方案爬过不少Cloudflare保护的页面,成功率基本能到90%以上。

3. 模拟真实用户行为,避免机械操作

就算隐藏了WebDriver,过于规律的操作还是会被盯上,建议补充这些细节:

  • 每次请求前随机等待1-3秒(用time.sleep(random.uniform(1,3)))
  • 进入页面后模拟鼠标滚动(比如driver.execute_script("window.scrollTo(0, document.body.scrollHeight);"))
  • 不要刚加载完页面就立刻定位元素,给页面留2-3秒的缓冲时间

4. 备选方案:改用请求库+Cloudflare破解

如果Selenium的方案还是不行,可以试试用requests结合cfscrape库(专门破解Cloudflare的JS挑战),不过要注意页面是否是动态渲染的:

安装依赖:

pip install requests cfscrape

使用示例:

import cfscrape

scraper = cfscrape.create_scraper()
url = "https://warthunder.com/en/community/userinfo/?nick=Hunter_i86"
response = scraper.get(url)
# 这里可以解析response.text提取页面内容

不过这个方法只适用于静态加载的内容,如果玩家数据是通过JS动态渲染的,还是得回到Selenium的方案。

最后提一句,你计划的24小时/次的爬取频率非常合理,保持这个频率基本不会触发Gaijin的反爬限制,继续坚持就好。

内容的提问来源于stack exchange,提问作者Matthieu Raynaud de Fitte

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 05:23:18