如何使用scrapy-splash实现点击下页URL不变站点的分页爬取
该需求完全可以实现,你只需要修改Lua脚本支持模拟点击分页按钮,同时调整爬虫的页码循环逻辑即可,具体修改方案如下:
1 核心实现思路
- 目标站点分页为前端JS动态渲染,URL无变化,可通过Splash在Lua脚本中模拟点击下一页按钮,等待页面加载完成后返回新的页面源码
- 给Lua脚本传递当前需要爬取的页码参数,循环点击对应次数的下一页按钮
- 增加下一页按钮可点击状态判断,避免无意义的重复点击或等待
2 完整修改后的代码
首先修正你原有代码的缩进问题(原代码中start_requests、parse方法没有写在类内部,属于语法错误),再替换对应逻辑即可:
import scrapy from scrapy_splash import SplashRequest from coins.items import CoinsItem class CoinsSpiderSpider(scrapy.Spider): name = 'coins_spider' allowed_domains = ['livecoinwatch.com'] start_urls = ['https://www.livecoinwatch.com'] # 自定义要爬取的总页数 MAX_PAGES = 3 lua_script = ''' function main(splash, args) splash.private_mode_enabled = false url = args.url -- 接收外部传入的目标页码 target_page = args.page or 1 headers = { ['User-Agent'] = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/94.0.4606.71 Safari/537.36 Edg/94.0.992.38' } splash:set_custom_headers(headers) assert(splash:go(url)) assert(splash:wait(3)) splash:set_viewport_full() -- 非第一页的情况,循环点击下一页 if target_page > 1 then for i=1, target_page-1 do -- 定位下一页按钮 local next_btn = splash:select('button.pagination-item.next') -- 按钮不可点击说明已经到最后一页,终止循环 if not next_btn or next_btn.attributes.disabled then break end next_btn:mouse_click() -- 等待表格数据加载完成 assert(splash:wait(2)) splash:wait_for_resume([[ function main(splash) { const check = setInterval(() => { const rows = document.querySelectorAll('tr.table-row.filter-row') if (rows.length >= 50) { clearInterval(check) splash.resume() } }, 500) } ]], 10) end end return splash:html() end ''' def start_requests(self): # 从第一页开始发起请求 for url in self.start_urls: yield SplashRequest( url=url, callback=self.parse, endpoint='execute', args={ 'lua_source': self.lua_script, 'page': 1 }, meta={'current_page': 1} ) def parse(self, response): current_page = response.meta.get('current_page', 1) rows = response.xpath('//tr[@class="table-row filter-row"]') for row in rows: item = CoinsItem() item['coin'] = row.xpath('./td[2]//div[@class="item-name ml10"]/div/text()').extract_first() item['price'] = row.xpath('./td[3]//text()').extract_first() item['marketCap'] = row.xpath('./td[4]/text()').extract_first() item['volumn24h'] = row.xpath('./td[5]/text()').extract_first() item['Liquidity'] = row.xpath('./td[6]/text()').extract_first() item['allTimeHigh'] = row.xpath('./td[7]/text()').extract_first() item['hour1_value'] = row.xpath('./td[8]/span/text()').extract_first() item['hour1_class'] = row.xpath('./td[8]/@class').extract_first() item['hour24_value'] = row.xpath('./td[9]/span/text()').extract_first() item['hour24_class'] = row.xpath('./td[9]/@class').extract_first() yield item # 分页逻辑:未到最大页数则继续爬下一页 if current_page < self.MAX_PAGES: next_page = current_page + 1 yield SplashRequest( url=response.url, callback=self.parse, endpoint='execute', args={ 'lua_source': self.lua_script, 'page': next_page }, meta={'current_page': next_page} )
3 注意事项
- 可根据自己的网络情况调整脚本中的等待时间,避免页面未完全加载就返回源码导致数据缺漏
- 建议在
settings.py中添加DOWNLOAD_DELAY配置设置请求间隔,避免请求频率过高被站点反爬拦截 - 若爬取过程中发现下一页按钮选择器失效,可通过浏览器F12重新定位下一页按钮的正确选择器替换即可
内容的提问来源于stack exchange,提问作者mxsmxm
相关产品推荐
相关产品推荐

