You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫求助:BeautifulSoup循环爬取MLB数据触发AttributeError

解决Python爬虫获取MLB比赛日期&球队时的AttributeError问题

嘿,我刚看了你的代码和问题,这个AttributeError其实很好解决——核心原因是你误解了页面元素的结构!

问题出在哪?

目标页面里,带有viCellBg1类的<td>元素分两种:一种是包含比赛日期的(里面有<span>标签),另一种是显示球队名称的(没有<span>)。你现在的循环是把所有viCellBg1的td都当成既有日期又有球队,所以当循环到球队行时,matchups.span会返回None,调用.text自然就报错了。

修复后的代码

我们可以给代码加个判断,区分日期行和球队行,安全地提取需要的内容:

from datetime import datetime
from flask import render_template
from testApp import app
from bs4 import BeautifulSoup
import requests

source = requests.get('http://www.vegasinsider.com/mlb/odds/las-vegas/').text
soup = BeautifulSoup(source, "lxml")
tbl = soup.find('table', class_='frodds-data-tbl')

# 初始化当前日期,用来关联后续的球队
current_game_date = ""
for matchup_cell in tbl.find_all('td', class_='viCellBg1'):
    # 先找日期span,找到的话更新当前日期
    date_element = matchup_cell.find('span')
    if date_element:
        current_game_date = date_element.text.strip()
        print(f"📅 比赛日期: {current_game_date}")
    else:
        # 没有span就是球队行,提取球队名称
        team_element = matchup_cell.find('b').find('a')
        if team_element:
            team_name = team_element.text.strip()
            print(f"🏟️ 球队: {team_name}")
    print()

为什么这样改?

  1. 元素区分逻辑:通过find('span')判断当前单元格是日期行还是球队行,避免强行访问不存在的属性。
  2. 安全提取:提取球队时先确认<b>和<a>标签存在,防止再次出现NoneType错误。
  3. 关联日期与球队:用current_game_date变量保存当前日期,后续的球队都会对应这个日期,符合页面的排版逻辑。

这样调整后,你的代码就能正常输出所有比赛日期和对应的球队列表啦!

内容的提问来源于stack exchange,提问作者kenneth2k1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 10:00:03