Python网页爬虫报错IndexError: list index out of range如何修复
足球赔率爬虫IndexError问题修复方案
问题根因
你遇到的IndexError: list index out of range核心原因是:pandas读取的表格行数与BeautifulSoup抓取的比分列表行数不一致,当遍历df的行索引number超过scores列表的最大索引时,就会触发越界。
触发行数不一致的常见场景:
- 页面未完全加载就抓取源码,部分行还没渲染出来
pd.read_html会自动过滤部分空行、格式异常行,但是BS的选择器会保留所有匹配的行,反之亦然- 你在
parse_data中重复读取browser.page_source覆盖了传入的html参数,若两次读取间隔页面发生动态更新,就会出现数据源偏差
排查步骤
在parse_data函数中,拿到df和scores后立刻打印两者长度,复现错误时即可验证是否为行数不一致问题:
df = pd.read_html(html, header=0)[0] soup = bs(html, "lxml") scores = [i.select_one('.table-score').text if i.select_one('.table-score') is not None else nan for i in soup.select('#table-matches tr:nth-of-type(n+2)')] # 新增排查代码 print(f"df行数:{len(df)}, scores长度:{len(scores)}")
修复方案
快速修复(兼容现有代码逻辑)
- 移除冗余的页面源码读取:删掉
parse_data函数里的html = browser.page_source行,统一使用传入的html参数,避免数据源不一致。 - 加索引边界判断,访问scores前先校验索引是否合法,越界时补空值:
# 替换原有的 game_data.score.append(scores[number]) if number < len(scores): game_data.score.append(scores[number]) else: game_data.score.append(nan)
- 加页面加载等待,避免未渲染完成就抓数据:
先导入依赖:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By
在browser.get(url)之后加显式等待:
WebDriverWait(browser, 10).until(EC.presence_of_element_located((By.ID, 'table-matches')))
最优修复(彻底解决对齐问题)
不要同时用pandas和BeautifulSoup两个数据源取字段,统一用BeautifulSoup遍历表格行,一次性取出所有需要的字段(时间、赛事、比分、赔率等),完全避免行索引不对齐的问题。
内容的提问来源于stack exchange,提问作者user16304089
相关产品推荐
相关产品推荐

