使用Scrapy Shell结合Splash抓取URL失败求助
解决Scrapy Splash因Recaptcha导致服务冻结的问题
针对你遇到的问题,以下是几个实用的解决方向:
给Splash请求添加明确超时参数
直接使用render.html时,Splash可能会无限等待页面加载(比如Recaptcha弹窗一直存在),加上wait和timeout参数强制终止请求:fetch('http://localhost:8050/render.html?url=https://www.barbiermotorsport.nl/motoren&wait=3&timeout=10')wait控制页面等待加载的最大时长(秒),timeout是Splash处理整个请求的超时阈值,避免单个请求拖垮服务。限制Splash的资源占用
如果用Docker运行Splash,默认资源配置可能不足以应对Recaptcha页面的渲染压力,启动时添加资源限制:docker run -p 8050:8050 --memory=2g --cpus=1 scrapinghub/splash这样即使页面渲染卡住,也不会耗尽主机资源导致Splash冻结。
用Lua脚本提前检测Recaptcha并终止渲染
自定义Splash的Lua脚本,在页面加载后检查是否存在Recaptcha元素,一旦检测到就立即返回结果,避免无限等待:function main(splash, args) splash:set_user_agent(args.ua) local ok, reason = splash:go(args.url) if not ok then return {error=reason} end -- 检测常见的Recaptcha选择器 local has_recaptcha = splash:select('#g-recaptcha') ~= nil or splash:select('.g-recaptcha') ~= nil if has_recaptcha then return { has_recaptcha = true, html = splash:html(), current_url = splash:url() } end -- 无验证码时再等待页面加载完成 splash:wait(args.wait or 2) return { html = splash:html(), current_url = splash:url() } end然后通过
execute接口调用这个脚本:fetch('http://localhost:8050/execute?url=https://www.barbiermotorsport.nl/motoren&lua_source=你的Lua脚本内容&ua=你的User-Agent')全局配置Splash超时
修改Splash的配置文件(Docker部署可通过挂载配置或环境变量),设置全局超时参数:
在splash.ini中添加:[splash] request_timeout = 10 render_timeout = 15这样所有请求都会受到超时限制,不会出现无限挂起的情况。
优先尝试静态抓取
对于目标站点,先尝试用普通Scrapy请求(不经过Splash)获取静态HTML,如果只有部分动态内容需要渲染,再针对性使用Splash,减少触发Recaptcha的概率。
内容的提问来源于stack exchange,提问作者sejteN
相关产品推荐
相关产品推荐

