You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取TOP10 ODI球员排名的Python爬虫代码运行报IndexError错误

报错产生原因
  • 本次IndexError: list index out of range的直接触发点是以下两行代码:
player = row.find_all('td', {'class': 'table-body__cell u-center-text'})[0].text.strip()
team = row.find_all('td', {'class': 'table-body__cell u-center-text'})[1].text.strip()

find_all方法返回匹配到的节点列表,当列表长度小于2时,访问索引0或1都会触发越界报错。

  • 常见的具体诱因:
    1. 网页结构更新,或者类名拼写错误,导致table-body__cell u-center-text这个组合类名匹配到的td节点数量不足2个
    2. 遍历到的tr.table-body行存在非球员数据的无效行(比如分隔行、广告行、空白行),这类行没有对应结构的td节点
  • 额外隐藏错误:你的代码中存在变量名拼写错误,前面定义ranking为th节点对象,后续直接写ranking[player] = ...会抛出类型不支持索引的错误,你实际是要赋值到rankings字典。
解决方法
  1. 先修复变量名拼写错误,将ranking[player] = {'position':position,'player': player, 'team': team,'ranking': ranking}替换为rankings[player] = {'position':position,'player': player, 'team': team,'ranking': ranking}
  2. 对节点查找结果增加长度校验和存在性判断,避免直接硬取索引,修改后的遍历逻辑参考:
for row in table.find_all('tr', {'class': 'table-body'}):
    center_cells = row.find_all('td', {'class': 'table-body__cell u-center-text'})
    # 匹配到的节点不足2个就跳过当前无效行
    if len(center_cells) < 2:
        continue
    position_ele = row.find('td', {'class': 'table-body__cell table-body__cell--position u-text-right'})
    ranking_ele = row.find('td', {'class': 'table-body__cell u-text-right'})
    # 必要节点不存在就跳过
    if not position_ele or not ranking_ele:
        continue
    position = position_ele.text.strip()
    player = center_cells[0].text.strip()
    team = center_cells[1].text.strip()
    ranking = ranking_ele.text.strip()
    rankings[player] = {'position': position,'team': team,'ranking': ranking}
  1. 如果修改后仍然拿不到数据,可以打印当前行的html结构确认类名是否匹配,调整查找规则即可:print(row.prettify())

内容的提问来源于stack exchange,提问作者bishwajit bhattacharya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 20:57:02