You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy-Splash无法完成无限滚动,漏抓二手车数据求助

解决Scrapy-Splash爬取无限滚动页面漏抓问题的实用方案

嘿,我处理过好多类似的无限滚动爬取坑,结合你用Scrapy-Splash的场景,给你几个针对性的思路和代码示例,应该能帮你解决漏抓的问题:

1. 优化滚动Lua脚本,精准判断页面底部

很多时候漏抓是因为脚本没等页面加载完就停止滚动,或者没判断是否真的到了底部。试试这个循环滚动的脚本,它会对比滚动前后的页面高度,确认没有新内容加载后才停止:

function main(splash, args)
  splash:go(args.url)
  splash:wait(2) -- 初始等待页面核心内容加载
  
  local scroll_count = 0
  local max_scrolls = 25 -- 设个上限防止无限循环,可根据网站内容量调整
  local prev_scroll_height = 0

  while scroll_count < max_scrolls do
    -- 滚动到当前页面底部
    splash:runjs("window.scrollTo(0, document.body.scrollHeight);")
    -- 等待新内容加载,时间根据网站响应速度调整(比如图片多的话设长点)
    splash:wait(3)

    -- 获取当前页面总高度
    local curr_scroll_height = splash:evaljs("document.body.scrollHeight")
    -- 如果高度不再变化,说明已经加载完所有内容
    if curr_scroll_height == prev_scroll_height then
      break
    end
    prev_scroll_height = curr_scroll_height
    scroll_count = scroll_count + 1
  end

  return splash:html()
end

这个脚本的核心是用页面高度变化判断是否还有新内容,比单纯滚动固定次数靠谱得多。

2. 等待加载状态结束再继续滚动

有些网站加载新内容时会显示loading动画,这时候直接滚动可能会跳过内容。可以在脚本里加入等待loading元素消失的逻辑:

-- 在每次滚动后添加这段代码,假设loading元素的CSS选择器是'.loading-overlay'
local is_loading = splash:evaljs("document.querySelector('.loading-overlay') !== null")
while is_loading do
  splash:wait(1)
  is_loading = splash:evaljs("document.querySelector('.loading-overlay') !== null")
end

或者用Splash自带的等待选择器方法:

-- 等待loading元素消失,超时则继续(避免卡住)
splash:wait_for_selector(".loading-overlay", cancel_on_timeout=true)

3. 调整Splash的超时与等待配置

有时候页面加载慢,默认的等待时间不够导致内容没加载完。你可以在Scrapy的settings.py里调整这些参数:

# 延长Splash请求超时时间(单位:秒)
SPLASH_TIMEOUT = 60
# 调整Splash的默认等待时间
SPLASH_DEFAULT_WAIT = 3

同时在Lua脚本里,根据网站的实际加载速度灵活调整splash:wait()的时间,比如图片多的页面可以设为4-5秒。

4. 模拟更真实的用户滚动行为

有些网站会检测滚动的频率和幅度,太机械的滚动可能会被限制。试试分小幅度滚动+随机等待:

function main(splash, args)
  splash:go(args.url)
  splash:wait(2)

  local scroll_count = 0
  local max_scrolls = 30
  local prev_height = 0

  while scroll_count < max_scrolls do
    -- 每次滚动500px,模拟用户慢慢滑动
    splash:runjs("window.scrollBy(0, 500);")
    -- 随机等待1-3秒,更贴近真实用户行为
    splash:wait(math.random(1,3))

    local curr_height = splash:evaljs("document.body.scrollHeight")
    if curr_height == prev_height then
      break
    end
    prev_height = curr_height
    scroll_count = scroll_count + 1
  end

  return splash:html()
end

5. 排查Splash实例的状态问题

虽然你把并发降到1了,但有时候Splash的实例可能残留之前的会话状态,影响爬取。可以在请求Splash时加上magic_response=True,确保每次请求都用全新的上下文:

yield SplashRequest(
    url=url,
    callback=self.parse,
    endpoint='execute',
    args={
        'lua_source': your_lua_script,
        'magic_response': True  # 启用这个参数重置会话
    }
)

先从优化滚动脚本开始试,这是最常见的解决方向。如果还是有问题,可以打开Splash的调试模式(在请求里加debug=1),查看页面渲染后的HTML,确认是不是真的没加载到内容,还是解析环节出了问题。

内容的提问来源于stack exchange,提问作者nevster

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:38:18