You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬取Oddsortal网站报IndexError列表索引越界如何解决

修复思路

报错的核心原因是你遍历到的非分类行(不带dark类的tr)中存在无效行,提取到的td标签数量不足6,直接按下标访问就会触发索引越界。你可以根据需求选择以下两种修改方案:


方案1:跳过无效数据行(推荐)

绝大多数情况下td数量不足的行都是无意义的空行、广告行、占位行,直接跳过即可,不影响有效数据采集:

def generate_matches(table):
    tr_tags = table.findAll('tr')
    # 提前初始化避免首次遇到比赛行时country/league未定义报错
    country = ''
    league = ''
    for tr_tag in tr_tags:
        if 'class' in tr_tag.attrs and 'dark' in tr_tag['class']:
            th_tag = tr_tag.find('th', {'class': 'first2 tl'})
            a_tags = th_tag.findAll('a')
            country = a_tags[0].text
            league = a_tags[1].text
        else:
            td_tags = tr_tag.findAll('td')
            # 仅处理td数量≥6的有效行
            if len(td_tags) >= 6:
                yield [
                    td_tags[0].text, td_tags[1].text, td_tags[2].text, td_tags[3].text,
                    td_tags[4].text, td_tags[5].text, country, league
                ]

方案2:缺省值填充保留所有行

如果你需要保留所有行的记录,不足的字段可以用空值或者nan填充:

from math import nan

def generate_matches(table):
    tr_tags = table.findAll('tr')
    country = ''
    league = ''
    for tr_tag in tr_tags:
        if 'class' in tr_tag.attrs and 'dark' in tr_tag['class']:
            th_tag = tr_tag.find('th', {'class': 'first2 tl'})
            a_tags = th_tag.findAll('a')
            country = a_tags[0].text
            league = a_tags[1].text
        else:
            td_tags = tr_tag.findAll('td')
            # 按位置取前6个字段,不足补nan
            field_list = []
            for i in range(6):
                field_list.append(td_tags[i].text if i < len(td_tags) else nan)
            field_list.extend([country, league])
            yield field_list

额外优化

上述代码提前初始化了country和league变量,可以避免表格第一个行就是比赛数据时,变量未定义的报错问题。

内容的提问来源于stack exchange,提问作者leonardo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 12:57:01