Scrapy Splash无法执行浏览器可用JS滚动代码问题排查
问题分析与解决方案
一、滚动脚本不生效的原因及修复
你的Lua脚本存在两个核心问题,导致无法获取滚动加载的数据:
等待时间不足
你的JS循环执行5次滚动,每次间隔2秒,最后一次滚动在第8秒(4*2000)才触发,但你仅调用了splash:wait(5.0),滚动动作还未完成就已获取HTML,自然拿不到加载后的内容。setTimeout在Splash中的局限性
Splash的JavaScript运行环境与浏览器存在差异,setTimeout的回调可能无法被Splash的事件循环正确处理,导致滚动动作未实际执行。建议改用同步滚动+等待的方式,每次滚动后预留足够时间让数据加载。
修复后的Lua脚本
function main(splash, args) splash:go(args.url) splash:wait(5.0) -- 等待初始页面渲染完成 -- 同步循环执行滚动,替代setTimeout for i=1,5 do splash:runjs("window.scrollTo(0, document.body.scrollHeight - 1500)") splash:wait(2.5) -- 等待当前滚动触发的数据加载完成,可根据实际速度调整 end splash:wait(2.0) -- 最后预留时间确保所有内容加载完毕 return {html = splash:html()} end
- 简化冗余参数
你的SplashRequest重复传递了url参数,可简化为:yield SplashRequest(url, callback=self.parse, endpoint='execute', args={'lua_source': lua_script, 'viewport': '1920x1080'})
二、无需滚动直接获取全部数据的方案
模拟滚动效率低且不稳定,更优方案是直接调用网站的AJAX接口获取数据:
抓包分析接口
在浏览器开发者工具的「网络」标签下,滚动页面时会捕获到类似/api/v2/search/...的请求,这类接口就是加载更多餐厅数据的核心接口,参数通常包含分页信息(page、pageSize)、地区ID等。直接请求API示例
找到接口后,用Scrapy直接请求并解析JSON数据,无需依赖Splash:import scrapy import json class OpenriceTstSpider(scrapy.Spider): name = "openrice_tst" allowed_domains = ["www.openrice.com"] # 替换为抓包得到的实际API地址 api_url = "https://www.openrice.com/api/v2/search/restaurants?districtId=...&page={}&pageSize=20" start_page = 1 def start_requests(self): yield scrapy.Request(self.api_url.format(self.start_page), callback=self.parse_api) def parse_api(self, response): data = json.loads(response.text) # 解析餐厅数据 for restaurant in data.get('restaurants', []): item = {} item['name'] = restaurant.get('name') item['rating'] = restaurant.get('rating') # 补充其他需要的字段 yield item # 分页逻辑:判断是否还有下一页 total_pages = data.get('totalPages') current_page = data.get('currentPage') if current_page < total_pages: next_page = current_page + 1 yield scrapy.Request(self.api_url.format(next_page), callback=self.parse_api)
这种方式效率更高,且能避免模拟滚动带来的加载延迟、环境差异等问题。
内容的提问来源于stack exchange,提问作者阿聰MrOnion
相关产品推荐
相关产品推荐

