Scrapy Splash:如何用Lua脚本实现非硬编码的分页循环
优化Splash Lua脚本:自动检测并点击分页下一页
问题背景
现有Lua脚本通过硬编码for i=1,4,1 do控制分页点击次数,但不同目标链接的分页数量不同,无法通用。需要改为自动检测下一页按钮是否存在,存在则点击,直到无下一页为止,此前尝试while循环时出现报错。
修改后的脚本代码
function main(splash, args) local my_user_agent = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/103.0.0.0 Safari/537.36' splash:set_user_agent(my_user_agent) splash.images_enabled = false assert(splash:go(args.url)) assert(splash:wait(0.5)) splash:set_viewport_full() local treat = require('treat') local result = {} -- 加入初始页面的HTML(若无需初始页可删除此行) table.insert(result, splash:html()) while true do -- 检测下一页按钮是否存在:适配Oddsportal的分页结构 local has_next_page = splash:runjs([[ const nextBtn = document.querySelector('#pagination a.next') || document.querySelector('#pagination a:contains("»")'); Boolean(nextBtn); ]]) if not has_next_page then break -- 无下一页时退出循环 end -- 执行点击操作,用容错逻辑避免脚本崩溃 local click_ok = splash:runjs([[ const nextBtn = document.querySelector('#pagination a.next') || document.querySelector('#pagination a:contains("»")'); if (nextBtn) { nextBtn.click(); true; } else { false; } ]]) if not click_ok then break end -- 等待页面加载完成 assert(splash:wait(2)) -- 将当前页面HTML加入结果集 table.insert(result, splash:html()) end return { current_url = splash:url(), pages_html = treat.as_array(result), screenshot = splash:png() } end
关键优化点
- 动态下一页检测:放弃原脚本中固定位置的
nth-child(7)选择器,改用Oddsportal分页的通用标识(a.next类或包含»的按钮),适配不同页数的分页结构,避免因页面布局变化导致元素定位失败。 - 循环逻辑重构:用
while true循环配合存在性检测实现自动终止,无需硬编码点击次数,完全适配任意页数的目标链接。 - 容错机制:通过
runjs返回的布尔值判断点击是否成功,避免因元素加载延迟或意外情况导致脚本直接终止。 - 结果集管理:用
table.insert自动维护结果数组,无需手动管理索引,代码更简洁。
内容的提问来源于stack exchange,提问作者benjamin olise
相关产品推荐
相关产品推荐

