You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy爬取adamchoi网站时生成空JSON文件问题求助

问题分析与解决方案

核心问题:Splash返回内容格式错误

你的Lua脚本返回的是包含PNG图片和HTML的数组,导致Scrapy收到的响应是JSON格式而非HTML页面。parse函数用XPath解析JSON字符串,自然无法匹配到任何表格行数据,最终生成空JSON文件。

其他潜在问题

  • 元素选择器可能失效:网站页面结构可能更新,原选择器label.btn.btn-sm.btn-primary可能无法定位到目标按钮
  • XPath选择器不够精确://tr会匹配所有表格行(包括表头),可能混入无效数据
  • 等待时间逻辑不够合理:页面加载和点击后的渲染需要针对性等待

修改后的代码

import scrapy
from scrapy_splash import SplashRequest

class AdamchoiSpider(scrapy.Spider):
    name = 'adamchoi'
    allowed_domains = ['www.adamchoi.co.uk']

    # 修正Lua脚本:仅返回HTML,调整元素选择器与等待逻辑
    script = '''
        function main(splash, args)
          splash.private_mode_enabled = false
          -- 访问目标页面
          assert(splash:go(args.url))
          -- 等待页面核心元素加载完成
          assert(splash:wait(2))
          -- 用更稳定的属性定位"All matches"按钮
          all_matches_btn = assert(splash:select("label[analytics-event='All matches']"))
          all_matches_btn:mouse_click()
          -- 等待点击后的数据渲染加载
          assert(splash:wait(3))
          splash:set_viewport_full()
          -- 仅返回页面HTML,供Scrapy直接解析
          return splash:html()
        end
    '''

    def start_requests(self):
        yield SplashRequest(
            url='https://www.adamchoi.co.uk/overs/detailed',
            callback=self.parse,
            endpoint='execute',
            args={'lua_source': self.script}
        )

    def parse(self, response):
        # 精确匹配数据行:只取目标表格tbody中的有效行
        rows = response.xpath('//table[@class="table-bordered"]/tbody/tr')

        for row in rows:
            date = row.xpath('./td[1]/text()').get()
            home_team = row.xpath('./td[2]/text()').get()
            score = row.xpath('./td[3]/text()').get()
            away_team = row.xpath('./td[4]/text()').get()
            # 过滤空数据,避免无效行混入结果
            if date and home_team and score and away_team:
                yield {
                    'date': date.strip(),
                    'home_team': home_team.strip(),
                    'score': score.strip(),
                    'away_team': away_team.strip(),
                }

额外检查项

  1. 确保Splash服务正常运行:本地需启动Splash容器(执行命令 docker run -p 8050:8050 scrapinghub/splash)
  2. 验证Scrapy配置:settings.py中需正确配置Splash相关中间件与地址
    DOWNLOADER_MIDDLEWARES = {
        'scrapy_splash.SplashCookiesMiddleware': 723,
        'scrapy_splash.SplashMiddleware': 725,
        'scrapy.downloadermiddlewares.httpcompression.HttpCompressionMiddleware': 810,
    }
    SPLASH_URL = 'http://localhost:8050'
    DUPEFILTER_CLASS = 'scrapy_splash.SplashAwareDupeFilter'
    HTTPCACHE_STORAGE = 'scrapy_splash.SplashAwareFSCacheStorage'
    
  3. 测试脚本有效性:可通过Splash Web界面(http://localhost:8050)测试Lua脚本,确认能正确渲染页面并点击按钮

内容的提问来源于stack exchange,提问作者DevSmartz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 19:43:13