You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用scrapy-splash实现点击下页URL不变站点的分页爬取

该需求完全可以实现,你只需要修改Lua脚本支持模拟点击分页按钮,同时调整爬虫的页码循环逻辑即可,具体修改方案如下:

1 核心实现思路
  • 目标站点分页为前端JS动态渲染,URL无变化,可通过Splash在Lua脚本中模拟点击下一页按钮,等待页面加载完成后返回新的页面源码
  • 给Lua脚本传递当前需要爬取的页码参数,循环点击对应次数的下一页按钮
  • 增加下一页按钮可点击状态判断,避免无意义的重复点击或等待
2 完整修改后的代码

首先修正你原有代码的缩进问题(原代码中start_requests、parse方法没有写在类内部,属于语法错误),再替换对应逻辑即可:

import scrapy
from scrapy_splash import SplashRequest
from coins.items import CoinsItem

class CoinsSpiderSpider(scrapy.Spider):
    name = 'coins_spider'
    allowed_domains = ['livecoinwatch.com']
    start_urls = ['https://www.livecoinwatch.com']
    # 自定义要爬取的总页数
    MAX_PAGES = 3

    lua_script = '''
        function main(splash, args)
            splash.private_mode_enabled = false
            url = args.url
            -- 接收外部传入的目标页码
            target_page = args.page or 1
            headers = {
                ['User-Agent'] = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/94.0.4606.71 Safari/537.36 Edg/94.0.992.38'
            }
            splash:set_custom_headers(headers)
            assert(splash:go(url)) 
            assert(splash:wait(3)) 
            splash:set_viewport_full()
            
            -- 非第一页的情况,循环点击下一页
            if target_page > 1 then
                for i=1, target_page-1 do
                    -- 定位下一页按钮
                    local next_btn = splash:select('button.pagination-item.next')
                    -- 按钮不可点击说明已经到最后一页,终止循环
                    if not next_btn or next_btn.attributes.disabled then
                        break
                    end
                    next_btn:mouse_click()
                    -- 等待表格数据加载完成
                    assert(splash:wait(2))
                    splash:wait_for_resume([[
                        function main(splash) {
                            const check = setInterval(() => {
                                const rows = document.querySelectorAll('tr.table-row.filter-row')
                                if (rows.length >= 50) {
                                    clearInterval(check)
                                    splash.resume()
                                }
                            }, 500)
                        }
                    ]], 10)
                end
            end
            return splash:html()
        end
    '''

    def start_requests(self):
        # 从第一页开始发起请求
        for url in self.start_urls:
            yield SplashRequest(
                url=url,
                callback=self.parse,
                endpoint='execute',
                args={
                    'lua_source': self.lua_script,
                    'page': 1
                },
                meta={'current_page': 1}
            )

    def parse(self, response):
        current_page = response.meta.get('current_page', 1)
        rows = response.xpath('//tr[@class="table-row filter-row"]')
        for row in rows:
            item = CoinsItem()
            item['coin'] = row.xpath('./td[2]//div[@class="item-name ml10"]/div/text()').extract_first()
            item['price'] = row.xpath('./td[3]//text()').extract_first()
            item['marketCap'] = row.xpath('./td[4]/text()').extract_first()
            item['volumn24h'] = row.xpath('./td[5]/text()').extract_first()
            item['Liquidity'] = row.xpath('./td[6]/text()').extract_first()
            item['allTimeHigh'] = row.xpath('./td[7]/text()').extract_first()
            item['hour1_value'] = row.xpath('./td[8]/span/text()').extract_first()
            item['hour1_class'] = row.xpath('./td[8]/@class').extract_first()
            item['hour24_value'] = row.xpath('./td[9]/span/text()').extract_first()
            item['hour24_class'] = row.xpath('./td[9]/@class').extract_first()
            yield item
        
        # 分页逻辑:未到最大页数则继续爬下一页
        if current_page < self.MAX_PAGES:
            next_page = current_page + 1
            yield SplashRequest(
                url=response.url,
                callback=self.parse,
                endpoint='execute',
                args={
                    'lua_source': self.lua_script,
                    'page': next_page
                },
                meta={'current_page': next_page}
            )
3 注意事项
  • 可根据自己的网络情况调整脚本中的等待时间,避免页面未完全加载就返回源码导致数据缺漏
  • 建议在settings.py中添加DOWNLOAD_DELAY配置设置请求间隔,避免请求频率过高被站点反爬拦截
  • 若爬取过程中发现下一页按钮选择器失效,可通过浏览器F12重新定位下一页按钮的正确选择器替换即可

内容的提问来源于stack exchange,提问作者mxsmxm

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 18:36:06