You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy Splash无法执行浏览器可用JS滚动代码问题排查

问题分析与解决方案

一、滚动脚本不生效的原因及修复

你的Lua脚本存在两个核心问题,导致无法获取滚动加载的数据:

  1. 等待时间不足
    你的JS循环执行5次滚动,每次间隔2秒,最后一次滚动在第8秒(4*2000)才触发,但你仅调用了splash:wait(5.0),滚动动作还未完成就已获取HTML,自然拿不到加载后的内容。

  2. setTimeout在Splash中的局限性
    Splash的JavaScript运行环境与浏览器存在差异,setTimeout的回调可能无法被Splash的事件循环正确处理,导致滚动动作未实际执行。建议改用同步滚动+等待的方式,每次滚动后预留足够时间让数据加载。

修复后的Lua脚本

function main(splash, args)
  splash:go(args.url)
  splash:wait(5.0) -- 等待初始页面渲染完成
  
  -- 同步循环执行滚动,替代setTimeout
  for i=1,5 do
    splash:runjs("window.scrollTo(0, document.body.scrollHeight - 1500)")
    splash:wait(2.5) -- 等待当前滚动触发的数据加载完成,可根据实际速度调整
  end
  
  splash:wait(2.0) -- 最后预留时间确保所有内容加载完毕
  return {html = splash:html()}
end
  1. 简化冗余参数
    你的SplashRequest重复传递了url参数,可简化为:
    yield SplashRequest(url, callback=self.parse, endpoint='execute', 
                        args={'lua_source': lua_script, 'viewport': '1920x1080'})
    

二、无需滚动直接获取全部数据的方案

模拟滚动效率低且不稳定,更优方案是直接调用网站的AJAX接口获取数据:

  1. 抓包分析接口
    在浏览器开发者工具的「网络」标签下,滚动页面时会捕获到类似/api/v2/search/...的请求,这类接口就是加载更多餐厅数据的核心接口,参数通常包含分页信息(page、pageSize)、地区ID等。

  2. 直接请求API示例
    找到接口后,用Scrapy直接请求并解析JSON数据,无需依赖Splash:

    import scrapy
    import json
    
    class OpenriceTstSpider(scrapy.Spider):
        name = "openrice_tst"
        allowed_domains = ["www.openrice.com"]
        # 替换为抓包得到的实际API地址
        api_url = "https://www.openrice.com/api/v2/search/restaurants?districtId=...&page={}&pageSize=20"
        start_page = 1
    
        def start_requests(self):
            yield scrapy.Request(self.api_url.format(self.start_page), callback=self.parse_api)
    
        def parse_api(self, response):
            data = json.loads(response.text)
            # 解析餐厅数据
            for restaurant in data.get('restaurants', []):
                item = {}
                item['name'] = restaurant.get('name')
                item['rating'] = restaurant.get('rating')
                # 补充其他需要的字段
                yield item
            
            # 分页逻辑:判断是否还有下一页
            total_pages = data.get('totalPages')
            current_page = data.get('currentPage')
            if current_page < total_pages:
                next_page = current_page + 1
                yield scrapy.Request(self.api_url.format(next_page), callback=self.parse_api)
    

这种方式效率更高,且能避免模拟滚动带来的加载延迟、环境差异等问题。

内容的提问来源于stack exchange,提问作者阿聰MrOnion

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 01:15:05