Scrapy+Splash访问whoscored.com遭Incapsula即时拦截的原因及配置咨询
Scrapy + Splash 被Whoscored Incapsula即时拦截的问题分析与解决方案
问题背景
你提到单独使用Splash访问whoscored.com时能成功请求2-3页才被Incapsula拦截,但结合Scrapy后首次请求就被拦截,返回以下拦截页面:
你的Scrapy配置如下:
BOT_NAME = 'scrapy_matchs' # Crawl responsibly by identifying yourself (and your website) on the user-agent #USER_AGENT = 'scrapy_matchs (+http://www.yourdomain.com)' # Obey robots.txt rules ROBOTSTXT_OBEY = False # Configure maximum concurrent requests performed by Scrapy (default: 16) CONCURRENT_REQUESTS = 1 # Configure a delay for requests for the same website (default: 0) # See https://docs.scrapy.org/en/latest/topics/settings.html#download-delay # See also autothrottle settings and docs DOWNLOAD_DELAY = 20 # The download delay setting will honor only one of: # CONCURRENT_REQUESTS_PER_DOMAIN = 1 CONCURRENT_REQUESTS_PER_IP = 1 # Disable cookies (enabled by default) COOKIES_ENABLED = False # Disable Telnet Console (enabled by default) #TELNETCONSOLE_ENABLED = False # Override the default request headers: DEFAULT_REQUEST_HEADERS = { 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8', 'Accept-Charset': 'ISO-8859-1,utf-8;q=0.7,*;q=0.3', 'Accept-Encoding': 'none', 'Accept-Language': 'en-US,en;q=0.8', 'Connection': 'keep-alive' } # Enable or disable spider middlewares # See https://docs.scrapy.org/en/latest/topics/spider-middleware.html SPIDER_MIDDLEWARES = { 'scrapy_splash.SplashDeduplicateArgsMiddleware': 100, } # Enable or disable downloader middlewares # See https://docs.scrapy.org/en/latest/topics/downloader-middleware.html DOWNLOADER_MIDDLEWARES = { 'scrapy.downloadermiddlewares.useragent.UserAgentMiddleware': None, 'scrapy_useragents.downloadermiddlewares.useragents.UserAgentsMiddleware': 500, 'scrapy_splash.SplashCookiesMiddleware': 723, 'scrapy_splash.SplashMiddleware': 725, 'scrapy.downloadermiddlewares.httpcompression.HttpCompressionMiddleware': 810, } USER_AGENTS = [ ('Mozilla/5.0 (X11; Linux x86_64) ' 'AppleWebKit/537.36 (KHTML, like Gecko) ' 'Chrome/57.0.2987.110 ' 'Safari/537.36'), # chrome ('Mozilla/5.0 (X11; Linux x86_64) ' 'AppleWebKit/537.36 (KHTML, like Gecko) ' 'Chrome/61.0.3163.79 ' 'Safari/537.36'), # chrome ('Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:55.0) ' 'Gecko/20100101 ' 'Firefox/55.0'), # firefox ('Mozilla/5.0 (X11; Linux x86_64) ' 'AppleWebKit/537.36 (KHTML, like Gecko) ' 'Chrome/61.0.3163.91 ' 'Safari/537.36'), # chrome ('Mozilla/5.0 (X11; Linux x86_64) ' 'AppleWebKit/537.36 (KHTML, like Gecko) ' 'Chrome/62.0.3202.89 ' 'Safari/537.36'), # chrome ('Mozilla/5.0 (X11; Linux x86_64) ' 'AppleWebKit/537.36 (KHTML, like Gecko) ' 'Chrome/63.0.3239.108 ' 'Safari/537.36'), # chrome ] # Enable or disable extensions # See https://docs.scrapy.org/en/latest/topics/extensions.html #EXTENSIONS = { # 'scrapy.extensions.telnet.TelnetConsole': None, #} # Configure item pipelines # See https://docs.scrapy.org/en/latest/topics/item-pipeline.html #ITEM_PIPELINES = { # 'scrapy_matchs.pipelines.ScrapyMatchsPipeline': 300, #} # Enable and configure the AutoThrottle extension (disabled by default) # See https://docs.scrapy.org/en/latest/topics/autothrottle.html AUTOTHROTTLE_ENABLED = True # The initial download delay AUTOTHROTTLE_START_DELAY = 30 # The maximum download delay to be set in case of high latencies AUTOTHROTTLE_MAX_DELAY = 60 # The average number of requests Scrapy should be sending in parallel to # each remote server AUTOTHROTTLE_TARGET_CONCURRENCY = 1.0 # Enable showing throttling stats for every response received: # AUTOTHROTTLE_DEBUG = False # Enable and configure HTTP caching (disabled by default) # See https://docs.scrapy.org/en/latest/topics/downloader-middleware.html#httpcache-middleware-settings #HTTPCACHE_ENABLED = True #HTTPCACHE_EXPIRATION_SECS = 0 #HTTPCACHE_DIR = 'httpcache' #HTTPCACHE_IGNORE_HTTP_CODES = [] #HTTPCACHE_STORAGE = 'scrapy.extensions.httpcache.FilesystemCacheStorage' SPLASH_URL = 'http://localhost:8050/' DUPEFILTER_CLASS = 'scrapy_splash.SplashAwareDupeFilter' HTTPCACHE_STORAGE = 'scrapy_splash.SplashAwareFSCacheStorage'
核心原因分析
我之前处理过不少Incapsula反爬的案例,这种差异主要是因为Scrapy + Splash组合的请求特征和单独用Splash时完全不一样,Incapsula的反爬系统对这些细节非常敏感:
- Cookie会话被打断:你设置了
COOKIES_ENABLED = False,但Incapsula依赖会话Cookie来验证“真实用户”的身份——单独用Splash时,它会自动维护浏览器级别的Cookie会话,但Scrapy禁用Cookie后,这个会话链就断了,直接触发拦截。 - 请求头不够“真实”:虽然你配了随机UA,但Scrapy默认的请求头缺失了一些真实浏览器必带的字段(比如
Referer、DNT),而且Splash渲染时的UA可能没有被正确传递给目标网站。 - Splash渲染太“机械”:单独用Splash时你可能手动调过等待时间、图片加载这些参数,但结合Scrapy后如果用了默认配置,渲染的页面没有动态交互(比如滚动),很容易被识别为爬虫。
- 延迟配置的冲突:你同时开了固定延迟和AutoThrottle,Scrapy会优先用AutoThrottle,初始30秒延迟看似够,但Incapsula可能还会检测请求的时序规律,单一的延迟模式还是会被盯上。
具体解决方案
1. 先改核心配置,修复基础问题
首先调整几个关键配置项,解决Cookie和UA的核心问题:
# 必须启用Cookie,Incapsula靠这个识别合法会话 COOKIES_ENABLED = True # 把AutoThrottle的初始延迟拉到60秒,最大延迟120秒,进一步放慢请求节奏 AUTOTHROTTLE_START_DELAY = 60 AUTOTHROTTLE_MAX_DELAY = 120 # 给UA列表加几个Windows和macOS的,别全是Linux的,太扎眼 USER_AGENTS = [ # 保留你原来的Linux UA... ('Mozilla/5.0 (Windows NT 10.0; Win64; x64) ' 'AppleWebKit/537.36 (KHTML, like Gecko) ' 'Chrome/118.0.0.0 ' 'Safari/537.36'), ('Mozilla/5.0 (Macintosh; Intel Mac OS X 13_6) ' 'AppleWebKit/605.1.15 (KHTML, like Gecko) ' 'Version/16.6 ' 'Safari/605.1.15'), ]
2. 优化Splash请求参数,模拟真实用户
在你的Spider中使用SplashRequest而非普通Request,并配置更贴近真实用户的渲染参数:
import random from scrapy_splash import SplashRequest def start_requests(self): target_url = "https://whoscored.com/..." # 随机选择UA random_ua = random.choice(self.settings.get('USER_AGENTS')) yield SplashRequest( url=target_url, callback=self.parse, args={ 'wait': 5, # 等待5秒让JS完全加载 'images': 1, # 加载图片,模拟真实用户浏览 'render_all': True, # 渲染整个页面(包括动态加载的内容) 'user-agent': random_ua, }, headers={ 'Referer': 'https://www.google.com/', # 模拟从Google跳转过来 'DNT': '1', # 加入Do Not Track字段,更贴近真实浏览器 'Accept-Language': 'en-US,en;q=0.9', }, )
3. 用Splash Lua脚本模拟动态交互
通过自定义Lua脚本模拟滚动页面、等待特定元素加载,进一步降低被识别的概率:
def start_requests(self): lua_script = """ function main(splash, args) splash:set_user_agent(args.ua) splash:go(args.url) splash:wait(3) -- 模拟滚动页面到底部,触发动态内容加载 splash:runjs("window.scrollTo(0, document.body.scrollHeight);") splash:wait(2) -- 等待目标元素加载完成(比如比赛数据表格) splash:wait_for_element('.table-matches') return splash:html() end """ random_ua = random.choice(self.settings.get('USER_AGENTS')) yield SplashRequest( url="https://whoscored.com/...", callback=self.parse, endpoint='execute', args={ 'lua_source': lua_script, 'ua': random_ua, 'url': "https://whoscored.com/...", }, headers={ 'Referer': 'https://www.google.com/', } )
4. 额外建议
- 使用代理IP池:Incapsula会跟踪IP请求频率,使用代理池每请求更换一个IP能有效降低被封概率。
- 避免重复请求:确保
DUPEFILTER_CLASS = 'scrapy_splash.SplashAwareDupeFilter'正确配置,避免重复请求同一页面。 - 更新Splash版本:使用最新版的Splash,旧版本的JS渲染引擎可能存在兼容性问题,导致被反爬系统识别。
内容的提问来源于stack exchange,提问作者Jérémie Octeau
相关产品推荐
相关产品推荐

