You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Oddsportal日期循环爬虫触发IndexError问题求助

解决方案

1. 解决IndexError(pop空列表)

不用显式判断列表是否为空,直接用try-except捕获异常处理空列表的情况。当pop()触发IndexError时,跳过当前条目即可,完全符合你不想用a_tags = [] if span is None else span.find_all('a')这类判断的要求。

示例代码片段:

# 替换你触发报错的pop逻辑
a_tags = span.find_all('a')
try:
    target_link = a_tags.pop().get('href')
except IndexError:
    # 列表为空时,跳过当前赛事条目
    continue

2. 实现多日期循环爬取

Oddsportal页面顶部的日期导航栏包含「前一天」「后一天」的链接,通过定位该链接并循环请求,即可实现多日期遍历。结合日期范围控制,避免无限循环:

完整代码示例

import requests
from bs4 import BeautifulSoup
from datetime import datetime, timedelta

# 配置请求头,模拟浏览器
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
}

# 设置爬取的日期范围(示例:2024年1月1日到1月7日)
start_date = datetime(2024, 1, 1)
end_date = datetime(2024, 1, 7)
current_date = end_date
# 初始URL(以欧冠赛事为例,可替换为你需要的项目)
current_url = f'https://www.oddsportal.com/soccer/europe/champions-league/results/#/{current_date.strftime("%Y%m%d")}/'

while current_date >= start_date:
    # 请求当前日期页面
    response = requests.get(current_url, headers=headers)
    soup = BeautifulSoup(response.text, 'html.parser')
    
    # ----------------------
    # 这里放入你已实现的赛事数据收集逻辑
    # 示例(仅演示结构):
    matches = soup.select('.table-main tr.deactivate')
    for match in matches:
        span = match.select_one('.name > span')
        if not span:
            continue
        a_tags = span.find_all('a')
        try:
            match_link = a_tags.pop().get('href')
            # 执行你已有的数据收集操作...
        except IndexError:
            continue
    # ----------------------
    
    # 获取「前一天」的链接,继续循环
    try:
        prev_day_elem = soup.select_one('a[title="Previous day"]')
        prev_day_link = prev_day_elem['href']
        current_url = f'https://www.oddsportal.com{prev_day_link}'
        current_date -= timedelta(days=1)
    except (TypeError, KeyError):
        # 没有找到前一天的链接,退出循环
        break

关键说明

  • 异常处理:无论是pop()的IndexError,还是日期链接定位失败的TypeError/KeyError,都用try-except处理,完全规避显式的空值判断。
  • 日期控制:通过datetime模块设定爬取范围,每次循环后日期减1天,直到达到起始日期或无更多日期链接。
  • 页面跳转:支持直接按日期格式构造URL,或通过页面导航链接跳转,两种方式都能实现多日期遍历。

内容的提问来源于stack exchange,提问作者Rander

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 02:06:03