You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Basketball Reference的Team Misc表格提取数值数据?

提取Basketball Reference Team Misc表格纯数值数据的方法

步骤1:解析注释中的表格HTML

Basketball Reference会将部分表格内容包裹在HTML注释中,需要先从你获取的table_text里提取出注释内的表格代码:

import re
from bs4 import BeautifulSoup, Comment

# 从table_text中筛选包含目标表格的注释
comments = table_text.find_all(string=lambda text: isinstance(text, Comment))
for comment in comments:
    if 'id="team_misc"' in comment:
        table_soup = BeautifulSoup(comment, 'html.parser')
        break

步骤2:提取纯数值数据

定位到表格后,遍历数据行,筛选出仅包含数字的单元格内容:

# 获取表格主体区域
table_body = table_soup.find('tbody')
rows = table_body.find_all('tr')

numeric_data = []
for row in rows:
    # 跳过表头行和空行
    if row.get('class') and 'thead' in row.get('class'):
        continue
    # 获取当前行的所有单元格
    cells = row.find_all('td')
    row_numbers = []
    for cell in cells:
        text = cell.get_text(strip=True)
        # 匹配整数或浮点数格式的内容
        if re.match(r'^-?\d+(\.\d+)?$', text):
            # 区分整数和浮点数存储
            row_numbers.append(float(text) if '.' in text else int(text))
    if row_numbers:
        numeric_data.append(row_numbers)

# 输出提取到的纯数值列表
print(numeric_data)

关键说明

  • 必须处理HTML注释:网站为了渲染逻辑会把表格藏在注释里,直接解析table_text拿不到表格结构
  • 正则匹配确保只提取纯数值:过滤掉文字、排名符号(如*)等非数值内容
  • 跳过冗余行:表格里的表头行(如Team、Lg Rank对应的行)需要跳过,只保留数据行

内容的提问来源于stack exchange,提问作者apersoninneed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 11:05:17