You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup抓取NFL数据时如何跳过重复表头行?

解决BeautifulSoup抓取NFL梦幻橄榄球表格的NoneType属性错误

问题原因

你遇到的'NoneType' object has no attribute 'a'错误,是因为目标表格每隔30行就会插入一行带thead类的重复表头,这类行里没有普通数据行的<a>标签,导致row.find('td', attrs={'data-stat': 'player'})返回None,后续调用dat.a自然会报错。

修复方案

直接在遍历行时过滤掉带thead类的行,或者先判断dat是否有效再处理。下面是两种可行的修改代码:

方法1:筛选有效行(推荐)

直接只获取不带thead类的<tr>元素,从根源避免无效行:

import requests
from bs4 import BeautifulSoup

url = 'https://www.pro-football-reference.com'
year = 2018

r = requests.get(url + '/years/' + str(year) + '/fantasy.htm')
soup = BeautifulSoup(r.content, 'html.parser')
parsed_table = soup.find_all('table')[0]  

# 筛选出不带thead类的行,跳过前2行表头
for i, row in enumerate(parsed_table.find_all('tr', class_=lambda c: c != 'thead')[2:]):
    print(i)
    dat = row.find('td', attrs={'data-stat': 'player'})
    if dat and dat.a:  # 双重保险,避免意外情况
        name = dat.a.get_text()
        stub = dat.a.get('href')
        # 这里可以添加处理数据的逻辑,比如存储name和stub

方法2:遍历中跳过无效行

如果不想修改行的筛选逻辑,也可以在循环里判断并跳过:

import requests
from bs4 import BeautifulSoup

url = 'https://www.pro-football-reference.com'
year = 2018

r = requests.get(url + '/years/' + str(year) + '/fantasy.htm')
soup = BeautifulSoup(r.content, 'html.parser')
parsed_table = soup.find_all('table')[0]  

for i, row in enumerate(parsed_table.find_all('tr')[2:]):
    # 检查当前行是否是thead行,是则跳过
    if 'thead' in row.get('class', []):
        continue
    print(i)
    dat = row.find('td', attrs={'data-stat': 'player'})
    if dat and dat.a:
        name = dat.a.get_text()
        stub = dat.a.get('href')
        # 处理数据

说明

  • 两种方法都加入了if dat and dat.a的判断,是为了应对表格中可能存在的其他异常行(比如空行),让代码更健壮。
  • 原代码遗漏了import requests,修复时已经补上,否则会报错。

内容的提问来源于stack exchange,提问作者codextrmz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 13:06:24