You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium爬取futbol24.com遭Bot识别,求解决方案

针对futbol24.com反爬的解决思路
  • 消除chromedriver的自动化标识
    网站通常会通过navigator.webdriver属性识别自动化工具,直接禁用这个特征:

    from selenium import webdriver
    from selenium.webdriver.chrome.options import Options
    
    options = Options()
    # 禁用自动化控制特征
    options.add_argument("--disable-blink-features=AutomationControlled")
    driver = webdriver.Chrome(options=options)
    # 额外注入JS覆盖webdriver属性
    driver.execute_script("Object.defineProperty(navigator, 'webdriver', {get: () => undefined})")
    
  • 模拟真实用户的交互行为
    纯静态加载页面很容易被识别,加入随机延迟和基础交互:

    import time
    import random
    
    # 页面加载后随机等待2-5秒
    time.sleep(random.uniform(2, 5))
    # 模拟滚动页面
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight/2);")
    time.sleep(random.uniform(1, 3))
    
  • 复用本地Chrome的用户数据
    直接调用你日常使用的Chrome配置,让爬取环境和手动访问完全一致:

    # 替换成你本地Chrome的用户数据目录路径
    options.add_argument(r"user-data-dir=C:\Users\你的用户名\AppData\Local\Google\Chrome\User Data")
    # 注意:启动前必须关闭所有本地Chrome窗口,否则会进程冲突
    
  • 严格控制请求频率
    短时间内高频请求是Bot的典型特征,每次请求间隔设置为3-10秒的随机值,完全模拟人类浏览节奏。

  • 对齐真实请求头信息
    除了User-Agent,Accept-Language、Accept-Encoding等字段也可能被检测。打开Chrome开发者工具的Network面板,查看手动访问时的请求头,然后在代码中添加:

    options.add_argument("accept-language=en-US,en;q=0.9")
    options.add_argument("accept-encoding=gzip, deflate, br")
    

内容的提问来源于stack exchange,提问作者Sav

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 15:40:27