You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy+Splash访问whoscored.com遭Incapsula即时拦截的原因及配置咨询

Scrapy + Splash 被Whoscored Incapsula即时拦截的问题分析与解决方案

问题背景

你提到单独使用Splash访问whoscored.com时能成功请求2-3页才被Incapsula拦截,但结合Scrapy后首次请求就被拦截,返回以下拦截页面:

你的Scrapy配置如下:

BOT_NAME = 'scrapy_matchs'
# Crawl responsibly by identifying yourself (and your website) on the user-agent
#USER_AGENT = 'scrapy_matchs (+http://www.yourdomain.com)'
# Obey robots.txt rules
ROBOTSTXT_OBEY = False
# Configure maximum concurrent requests performed by Scrapy (default: 16)
CONCURRENT_REQUESTS = 1
# Configure a delay for requests for the same website (default: 0)
# See https://docs.scrapy.org/en/latest/topics/settings.html#download-delay
# See also autothrottle settings and docs
DOWNLOAD_DELAY = 20
# The download delay setting will honor only one of:
# CONCURRENT_REQUESTS_PER_DOMAIN = 1
CONCURRENT_REQUESTS_PER_IP = 1
# Disable cookies (enabled by default)
COOKIES_ENABLED = False
# Disable Telnet Console (enabled by default)
#TELNETCONSOLE_ENABLED = False
# Override the default request headers:
DEFAULT_REQUEST_HEADERS = {
 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8',
 'Accept-Charset': 'ISO-8859-1,utf-8;q=0.7,*;q=0.3',
 'Accept-Encoding': 'none',
 'Accept-Language': 'en-US,en;q=0.8',
 'Connection': 'keep-alive'
}
# Enable or disable spider middlewares
# See https://docs.scrapy.org/en/latest/topics/spider-middleware.html
SPIDER_MIDDLEWARES = {
 'scrapy_splash.SplashDeduplicateArgsMiddleware': 100,
}
# Enable or disable downloader middlewares
# See https://docs.scrapy.org/en/latest/topics/downloader-middleware.html
DOWNLOADER_MIDDLEWARES = {
 'scrapy.downloadermiddlewares.useragent.UserAgentMiddleware': None,
 'scrapy_useragents.downloadermiddlewares.useragents.UserAgentsMiddleware': 500,
 'scrapy_splash.SplashCookiesMiddleware': 723,
 'scrapy_splash.SplashMiddleware': 725,
 'scrapy.downloadermiddlewares.httpcompression.HttpCompressionMiddleware': 810,
}
USER_AGENTS = [
 ('Mozilla/5.0 (X11; Linux x86_64) '
 'AppleWebKit/537.36 (KHTML, like Gecko) '
 'Chrome/57.0.2987.110 '
 'Safari/537.36'), # chrome
 ('Mozilla/5.0 (X11; Linux x86_64) '
 'AppleWebKit/537.36 (KHTML, like Gecko) '
 'Chrome/61.0.3163.79 '
 'Safari/537.36'), # chrome
 ('Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:55.0) '
 'Gecko/20100101 '
 'Firefox/55.0'), # firefox
 ('Mozilla/5.0 (X11; Linux x86_64) '
 'AppleWebKit/537.36 (KHTML, like Gecko) '
 'Chrome/61.0.3163.91 '
 'Safari/537.36'), # chrome
 ('Mozilla/5.0 (X11; Linux x86_64) '
 'AppleWebKit/537.36 (KHTML, like Gecko) '
 'Chrome/62.0.3202.89 '
 'Safari/537.36'), # chrome
 ('Mozilla/5.0 (X11; Linux x86_64) '
 'AppleWebKit/537.36 (KHTML, like Gecko) '
 'Chrome/63.0.3239.108 '
 'Safari/537.36'), # chrome
]
# Enable or disable extensions
# See https://docs.scrapy.org/en/latest/topics/extensions.html
#EXTENSIONS = {
# 'scrapy.extensions.telnet.TelnetConsole': None,
#}
# Configure item pipelines
# See https://docs.scrapy.org/en/latest/topics/item-pipeline.html
#ITEM_PIPELINES = {
# 'scrapy_matchs.pipelines.ScrapyMatchsPipeline': 300,
#}
# Enable and configure the AutoThrottle extension (disabled by default)
# See https://docs.scrapy.org/en/latest/topics/autothrottle.html
AUTOTHROTTLE_ENABLED = True
# The initial download delay
AUTOTHROTTLE_START_DELAY = 30
# The maximum download delay to be set in case of high latencies
AUTOTHROTTLE_MAX_DELAY = 60
# The average number of requests Scrapy should be sending in parallel to
# each remote server
AUTOTHROTTLE_TARGET_CONCURRENCY = 1.0
# Enable showing throttling stats for every response received:
# AUTOTHROTTLE_DEBUG = False
# Enable and configure HTTP caching (disabled by default)
# See https://docs.scrapy.org/en/latest/topics/downloader-middleware.html#httpcache-middleware-settings
#HTTPCACHE_ENABLED = True
#HTTPCACHE_EXPIRATION_SECS = 0
#HTTPCACHE_DIR = 'httpcache'
#HTTPCACHE_IGNORE_HTTP_CODES = []
#HTTPCACHE_STORAGE = 'scrapy.extensions.httpcache.FilesystemCacheStorage'
SPLASH_URL = 'http://localhost:8050/'
DUPEFILTER_CLASS = 'scrapy_splash.SplashAwareDupeFilter'
HTTPCACHE_STORAGE = 'scrapy_splash.SplashAwareFSCacheStorage'

核心原因分析

我之前处理过不少Incapsula反爬的案例,这种差异主要是因为Scrapy + Splash组合的请求特征和单独用Splash时完全不一样,Incapsula的反爬系统对这些细节非常敏感:

  1. Cookie会话被打断:你设置了COOKIES_ENABLED = False,但Incapsula依赖会话Cookie来验证“真实用户”的身份——单独用Splash时,它会自动维护浏览器级别的Cookie会话,但Scrapy禁用Cookie后,这个会话链就断了,直接触发拦截。
  2. 请求头不够“真实”:虽然你配了随机UA,但Scrapy默认的请求头缺失了一些真实浏览器必带的字段(比如Referer、DNT),而且Splash渲染时的UA可能没有被正确传递给目标网站。
  3. Splash渲染太“机械”:单独用Splash时你可能手动调过等待时间、图片加载这些参数,但结合Scrapy后如果用了默认配置,渲染的页面没有动态交互(比如滚动),很容易被识别为爬虫。
  4. 延迟配置的冲突:你同时开了固定延迟和AutoThrottle,Scrapy会优先用AutoThrottle,初始30秒延迟看似够,但Incapsula可能还会检测请求的时序规律,单一的延迟模式还是会被盯上。

具体解决方案

1. 先改核心配置,修复基础问题

首先调整几个关键配置项,解决Cookie和UA的核心问题:

# 必须启用Cookie,Incapsula靠这个识别合法会话
COOKIES_ENABLED = True

# 把AutoThrottle的初始延迟拉到60秒,最大延迟120秒,进一步放慢请求节奏
AUTOTHROTTLE_START_DELAY = 60
AUTOTHROTTLE_MAX_DELAY = 120

# 给UA列表加几个Windows和macOS的,别全是Linux的,太扎眼
USER_AGENTS = [
    # 保留你原来的Linux UA...
    ('Mozilla/5.0 (Windows NT 10.0; Win64; x64) '
     'AppleWebKit/537.36 (KHTML, like Gecko) '
     'Chrome/118.0.0.0 '
     'Safari/537.36'),
    ('Mozilla/5.0 (Macintosh; Intel Mac OS X 13_6) '
     'AppleWebKit/605.1.15 (KHTML, like Gecko) '
     'Version/16.6 '
     'Safari/605.1.15'),
]

2. 优化Splash请求参数,模拟真实用户

在你的Spider中使用SplashRequest而非普通Request,并配置更贴近真实用户的渲染参数:

import random
from scrapy_splash import SplashRequest

def start_requests(self):
    target_url = "https://whoscored.com/..."
    # 随机选择UA
    random_ua = random.choice(self.settings.get('USER_AGENTS'))
    yield SplashRequest(
        url=target_url,
        callback=self.parse,
        args={
            'wait': 5,  # 等待5秒让JS完全加载
            'images': 1,  # 加载图片,模拟真实用户浏览
            'render_all': True,  # 渲染整个页面(包括动态加载的内容)
            'user-agent': random_ua,
        },
        headers={
            'Referer': 'https://www.google.com/',  # 模拟从Google跳转过来
            'DNT': '1',  # 加入Do Not Track字段,更贴近真实浏览器
            'Accept-Language': 'en-US,en;q=0.9',
        },
    )

3. 用Splash Lua脚本模拟动态交互

通过自定义Lua脚本模拟滚动页面、等待特定元素加载,进一步降低被识别的概率:

def start_requests(self):
    lua_script = """
    function main(splash, args)
        splash:set_user_agent(args.ua)
        splash:go(args.url)
        splash:wait(3)
        -- 模拟滚动页面到底部,触发动态内容加载
        splash:runjs("window.scrollTo(0, document.body.scrollHeight);")
        splash:wait(2)
        -- 等待目标元素加载完成(比如比赛数据表格)
        splash:wait_for_element('.table-matches')
        return splash:html()
    end
    """
    random_ua = random.choice(self.settings.get('USER_AGENTS'))
    yield SplashRequest(
        url="https://whoscored.com/...",
        callback=self.parse,
        endpoint='execute',
        args={
            'lua_source': lua_script,
            'ua': random_ua,
            'url': "https://whoscored.com/...",
        },
        headers={
            'Referer': 'https://www.google.com/',
        }
    )

4. 额外建议

  • 使用代理IP池:Incapsula会跟踪IP请求频率,使用代理池每请求更换一个IP能有效降低被封概率。
  • 避免重复请求:确保DUPEFILTER_CLASS = 'scrapy_splash.SplashAwareDupeFilter'正确配置,避免重复请求同一页面。
  • 更新Splash版本:使用最新版的Splash,旧版本的JS渲染引擎可能存在兼容性问题,导致被反爬系统识别。

内容的提问来源于stack exchange,提问作者Jérémie Octeau

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:37:36