You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取Basketball Reference时遇NoneType属性错误求助

问题描述

我正在为Basketball Reference网站开发数据爬虫,用于提取指定NBA球员“Per Game”表格的表头与统计数据,并展示在自建页面中。相关代码如下:

@app.route('/', methods=['GET', 'POST'])
def index():
  if request.method == 'POST':
    player_name = request.form['player']
    player_url = get_player_url(player_name)
    
    response = requests.get(player_url)
    soup = BeautifulSoup(response.content, 'html.parser')
    stats = soup.find('div', {'id':'all_per_game'})
    
    table = stats.find('table')
    table_headings = [th.get_text() for th in table.find('tr').find_all('th')]
    table_data = []
    for row in table.find_all('tr')[1:]:
      row_data = [td.get_text() for td in row.find_all('td')]
      table_data.append(row_data)
      
    return render_template('stats.html', player_name=player_name, headings=table_headings, data=table_data)
  else:
    return render_template('index.html', players=players)

运行本地Python应用时,在获取表格表头步骤出现如下错误:

File ¨main.py¨, line 48, in index
   table_headings = [th.get_text() for th in table.find('tr').find_all('th')]
AttributeError: 'NoneType' object has no attribute find

我是BeautifulSoup新手,已查阅网络资料但未找到适配方案,请问如何正确提取表头?该问题是否与table对象类型有关?


解决方案

这个错误的核心是table或table.find('tr')返回了None——也就是BeautifulSoup没找到对应的元素,大概率是Basketball Reference的页面结构有特殊处理,或是你的选择器/请求方式有问题。

问题排查与修复步骤

  1. 先确认元素是否存在,添加调试与异常处理
    先在代码中加入判断,避免None对象调用方法:

    stats = soup.find('div', {'id':'all_per_game'})
    # 先判断stats是否存在,再找表格
    table = stats.find('table') if stats else None
    # 如果table不存在,直接返回错误提示
    if not table:
        return render_template('error.html', message='无法找到目标数据表格')
    

    如果stats是None,要么是请求被网站拦截,要么是页面结构变更;如果stats存在但table是None,那大概率是表格被藏在了HTML注释里——Basketball Reference常用这种方式防爬虫。

  2. 处理藏在注释中的表格
    如果表格在注释内,需要先提取注释内容再解析:

    from bs4 import Comment
    
    stats = soup.find('div', {'id':'all_per_game'})
    table = None
    if stats:
        # 先尝试直接找表格
        table = stats.find('table')
        # 找不到的话遍历所有注释内容
        if not table:
            comments = stats.find_all(string=lambda text: isinstance(text, Comment))
            for comment in comments:
                comment_soup = BeautifulSoup(comment, 'html.parser')
                table = comment_soup.find('table')
                if table:
                    break
    
  3. 正确提取表头
    不要直接用table.find('tr'),优先找thead标签(标准表格结构中表头会放在这里),避免误取数据行:

    table_headings = []
    # 优先从thead提取表头
    thead = table.find('thead')
    if thead:
        header_row = thead.find('tr')
        table_headings = [th.get_text(strip=True) for th in header_row.find_all('th')]
    else:
        # 没有thead的话,取表格第一个tr作为表头
        header_row = table.find('tr')
        if header_row:
            table_headings = [th.get_text(strip=True) for th in header_row.find_all('th')]
    

    用strip=True可以自动去除文本前后的空格和换行,让表头更整洁。

  4. 添加请求头模拟浏览器
    很多网站会拦截无标识的爬虫请求,给requests.get加上浏览器请求头:

    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
    }
    response = requests.get(player_url, headers=headers)
    

完整修正后的核心代码片段

@app.route('/', methods=['GET', 'POST'])
def index():
    if request.method == 'POST':
        player_name = request.form['player']
        player_url = get_player_url(player_name)
        
        # 添加浏览器请求头
        headers = {
            'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
        }
        response = requests.get(player_url, headers=headers)
        soup = BeautifulSoup(response.content, 'html.parser')
        stats = soup.find('div', {'id':'all_per_game'})
        
        table = None
        if stats:
            table = stats.find('table')
            # 从注释中找表格
            if not table:
                from bs4 import Comment
                comments = stats.find_all(string=lambda text: isinstance(text, Comment))
                for comment in comments:
                    comment_soup = BeautifulSoup(comment, 'html.parser')
                    table = comment_soup.find('table')
                    if table:
                        break
        
        if not table:
            return render_template('error.html', message='无法找到球员的Per Game数据表格')
        
        # 提取表头
        table_headings = []
        thead = table.find('thead')
        if thead:
            header_row = thead.find('tr')
            table_headings = [th.get_text(strip=True) for th in header_row.find_all('th')]
        else:
            header_row = table.find('tr')
            if header_row:
                table_headings = [th.get_text(strip=True) for th in header_row.find_all('th')]
        
        # 提取表格数据
        table_data = []
        tbody = table.find('tbody')
        rows = tbody.find_all('tr') if tbody else table.find_all('tr')[1:]
        
        for row in rows:
            # 跳过重复的表头行
            if row.get('class') and 'thead' in row.get('class'):
                continue
            row_data = [td.get_text(strip=True) for td in row.find_all('td')]
            table_data.append(row_data)
        
        return render_template('stats.html', player_name=player_name, headings=table_headings, data=table_data)
    else:
        return render_template('index.html', players=players)

关键注意点

  • 必须加请求头,避免被网站反爬机制拦截。
  • Basketball Reference的表格常藏在HTML注释中,这是你之前找不到table的主要原因。
  • 提取表头优先用thead标签,比直接找tr更准确。
  • 全程添加异常判断,避免因页面结构变化导致程序崩溃。

内容的提问来源于stack exchange,提问作者SumeetS413

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 10:05:45