You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy Shell结合Splash抓取URL失败求助

解决Scrapy Splash因Recaptcha导致服务冻结的问题

针对你遇到的问题,以下是几个实用的解决方向:

  • 给Splash请求添加明确超时参数
    直接使用render.html时,Splash可能会无限等待页面加载(比如Recaptcha弹窗一直存在),加上wait和timeout参数强制终止请求:

    fetch('http://localhost:8050/render.html?url=https://www.barbiermotorsport.nl/motoren&wait=3&timeout=10')
    

    wait控制页面等待加载的最大时长(秒),timeout是Splash处理整个请求的超时阈值,避免单个请求拖垮服务。

  • 限制Splash的资源占用
    如果用Docker运行Splash,默认资源配置可能不足以应对Recaptcha页面的渲染压力,启动时添加资源限制:

    docker run -p 8050:8050 --memory=2g --cpus=1 scrapinghub/splash
    

    这样即使页面渲染卡住,也不会耗尽主机资源导致Splash冻结。

  • 用Lua脚本提前检测Recaptcha并终止渲染
    自定义Splash的Lua脚本,在页面加载后检查是否存在Recaptcha元素,一旦检测到就立即返回结果,避免无限等待:

    function main(splash, args)
      splash:set_user_agent(args.ua)
      local ok, reason = splash:go(args.url)
      if not ok then
        return {error=reason}
      end
      -- 检测常见的Recaptcha选择器
      local has_recaptcha = splash:select('#g-recaptcha') ~= nil or splash:select('.g-recaptcha') ~= nil
      if has_recaptcha then
        return {
          has_recaptcha = true,
          html = splash:html(),
          current_url = splash:url()
        }
      end
      -- 无验证码时再等待页面加载完成
      splash:wait(args.wait or 2)
      return {
        html = splash:html(),
        current_url = splash:url()
      }
    end
    

    然后通过execute接口调用这个脚本:

    fetch('http://localhost:8050/execute?url=https://www.barbiermotorsport.nl/motoren&lua_source=你的Lua脚本内容&ua=你的User-Agent')
    
  • 全局配置Splash超时
    修改Splash的配置文件(Docker部署可通过挂载配置或环境变量),设置全局超时参数:
    在splash.ini中添加:

    [splash]
    request_timeout = 10
    render_timeout = 15
    

    这样所有请求都会受到超时限制,不会出现无限挂起的情况。

  • 优先尝试静态抓取
    对于目标站点,先尝试用普通Scrapy请求(不经过Splash)获取静态HTML,如果只有部分动态内容需要渲染,再针对性使用Splash,减少触发Recaptcha的概率。

内容的提问来源于stack exchange,提问作者sejteN

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 04:45:08