You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬取OddsPortal赛事数据触发IndexError索引越界问题求解

报错根因
  • 你分别用pd.read_html和BeautifulSoup读取同一份表格的不同字段,两类工具对表格行的过滤规则不一致:read_html会自动跳过空行、分组表头行,而你构造scores列表时包含了这些行,两个列表的长度天然不匹配。
  • 循环中遇到非赛事行时你用了continue跳过,但enumerate生成的number索引仍会持续累加,最终索引值超过scores列表的总长度,触发越界错误。
改用Xpath提取的实现方案

你提供的是单条赛事比分的绝对路径Xpath,包含固定行索引tr[9]无法匹配所有行,需要改为通用的相对路径规则,同时建议所有字段都直接从行节点提取,从根源上避免索引对齐问题。

第一步:补充必要依赖导入

from selenium import webdriver
from selenium.webdriver.common.by import By
import pandas as pd
from math import nan
import time # 可选,加延时避免反爬

第二步:重写parse_data函数,移除原有bs4和pandas读表逻辑

def parse_data(browser):
    game_data = GameData()
    # 匹配表格内所有行,过滤广告行
    match_rows = browser.find_elements(By.XPATH, '//table[@id="table-matches"]/tbody/tr')
    current_country = ""
    current_league = ""

    for row in match_rows:
        row_class = row.get_attribute("class")
        # 匹配国家/联赛分组行
        if "dark" in row_class:
            try:
                league_info = row.find_element(By.XPATH, './th[@class="first2 tl"]/a[2]').text
                current_country, current_league = league_info.split(" » ", 1)
            except:
                continue
            continue
        # 匹配普通赛事行,提取字段
        try:
            match_time = row.find_element(By.XPATH, './td[1]').text
            if ":" not in match_time:
                continue
            match_name = row.find_element(By.XPATH, './td[2]').text
            # 提取比分,无比分时返回nan
            score_ele = row.find_elements(By.XPATH, './td[contains(@class, "table-score")]')
            match_score = score_ele[0].text if score_ele else nan
        except:
            continue
        # 写入数据
        game_data.country.append(current_country)
        game_data.league.append(current_league)
        game_data.game.append(match_name)
        game_data.score.append(match_score)
    return game_data

第三步:修改主函数调用逻辑

if __name__ == '__main__':
    start_url = "https://www.oddsportal.com/matches/soccer/"
    browser = webdriver.Chrome()
    results = None
    urls = get_urls(browser, start_url)
    urls.insert(0, start_url)

    for number, url in enumerate(urls):
        browser.get(url)
        time.sleep(2) # 加延时等待页面加载,避免反爬
        game_data = parse_data(browser)
        if game_data is None or not game_data.game:
            continue
        result = pd.DataFrame(game_data.__dict__)
        if results is None:
            results = result
        else:
            results = pd.concat([results, result], ignore_index=True) # append已弃用,改用concat
注意事项
  • 原代码中全局定义的browser和主函数里实例化的browser重复,建议删掉全局的browser = webdriver.Chrome()避免冲突。
  • pandas的append方法已在高版本废弃,改用pd.concat合并数据集。

内容的提问来源于stack exchange,提问作者user16304089

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 11:18:02