You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬取NFL赛事赔率表格无数据返回问题求助

问题分析与解决

你的Scrapy脚本无报错但无数据返回,核心问题出在XPath匹配逻辑和字符串转义上,具体问题点及修复方案如下:

1. HTML转义字符导致XPath匹配失败

代码里用了"(HTML转义的双引号),但Python字符串中不需要这种转义,XPath会把"当成字符串的一部分,导致无法匹配到class="soh1"的表格。

比如原代码中的:

response.xpath('//table[@class="soh1"]//tbody/tr')

要改成:

response.xpath('//table[@class="soh1"]//tbody/tr')

同理,week字段的XPath里的"也要替换成普通双引号。

2. 未限定目标表格范围

你需要爬取第4个表格开始的数据,但原XPath匹配了所有class="soh1"的表格,前3个非目标表格的结构和赛事表格不同,遍历这些表格的行时,会拿到大量空值,甚至没有有效行数据。

需要给XPath加上位置筛选,只匹配第4个及以后的表格:

//table[@class="soh1"][position() >= 4]//tbody/tr

3. Week字段的XPath逻辑错误

原代码中week的XPath是全局查找第4个表格的h3,这会导致所有行的week都是同一个值,且如果该位置没有h3,就会返回空。正确的做法是先遍历每个目标表格,获取该表格对应的周数,再遍历表格内的行。

修复后的完整代码

import scrapy


class NFLOddsSpider(scrapy.Spider):
    name = 'NFLOdds'
    allowed_domains = ['www.sportsoddshistory.com']
    start_urls = ['https://www.sportsoddshistory.com/nfl-game-season/?y=2022']

    def parse(self, response):
        # 遍历第4个及以后的class为soh1的表格
        for table in response.xpath('//table[@class="soh1"][position() >= 4]'):
            # 获取当前表格对应的周数(从表格内的h3提取)
            week = table.xpath('.//td/h3/text()').extract_first()
            # 遍历当前表格内的所有行
            for row in table.xpath('.//tbody/tr'):
                day = row.xpath('td[1]//text()').extract_first()
                date = row.xpath('td[2]//text()').extract_first()
                time = row.xpath('td[3]//text()').extract_first()
                AtFav = row.xpath('td[4]//text()').extract_first()
                favorite = row.xpath('td[5]//text()').extract_first()
                score = row.xpath('td[6]//text()').extract_first()
                spread = row.xpath('td[7]//text()').extract_first()
                AtDog = row.xpath('td[8]//text()').extract_first()
                underdog = row.xpath('td[9]//text()').extract_first()
                OvUn = row.xpath('td[10]//text()').extract_first()
                notes = row.xpath('td[11]//text()').extract_first()

                oddsTable = {
                    'day': day,
                    'date': date,
                    'time': time,
                    'AtFav': AtFav,
                    'favorite': favorite,
                    'score': score,
                    'spread': spread,
                    'AtDog': AtDog,
                    'underdog': underdog,
                    'OvUn': OvUn,
                    'notes': notes,
                    'week': week
                }
                yield oddsTable

额外说明

  • 修复后的代码先定位到每个目标表格,再获取对应周数,最后遍历行数据,确保每行都能关联正确的周信息。
  • 使用.开头的相对XPath(比如.//tbody/tr),确保只在当前表格范围内查找行,避免全局匹配导致的错误。

内容的提问来源于stack exchange,提问作者Leo Torres

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 00:40:08