如何从Basketball Reference的Team Misc表格提取数值数据?
提取Basketball Reference Team Misc表格纯数值数据的方法
步骤1:解析注释中的表格HTML
Basketball Reference会将部分表格内容包裹在HTML注释中,需要先从你获取的table_text里提取出注释内的表格代码:
import re from bs4 import BeautifulSoup, Comment # 从table_text中筛选包含目标表格的注释 comments = table_text.find_all(string=lambda text: isinstance(text, Comment)) for comment in comments: if 'id="team_misc"' in comment: table_soup = BeautifulSoup(comment, 'html.parser') break
步骤2:提取纯数值数据
定位到表格后,遍历数据行,筛选出仅包含数字的单元格内容:
# 获取表格主体区域 table_body = table_soup.find('tbody') rows = table_body.find_all('tr') numeric_data = [] for row in rows: # 跳过表头行和空行 if row.get('class') and 'thead' in row.get('class'): continue # 获取当前行的所有单元格 cells = row.find_all('td') row_numbers = [] for cell in cells: text = cell.get_text(strip=True) # 匹配整数或浮点数格式的内容 if re.match(r'^-?\d+(\.\d+)?$', text): # 区分整数和浮点数存储 row_numbers.append(float(text) if '.' in text else int(text)) if row_numbers: numeric_data.append(row_numbers) # 输出提取到的纯数值列表 print(numeric_data)
关键说明
- 必须处理HTML注释:网站为了渲染逻辑会把表格藏在注释里,直接解析
table_text拿不到表格结构 - 正则匹配确保只提取纯数值:过滤掉文字、排名符号(如
*)等非数值内容 - 跳过冗余行:表格里的表头行(如Team、Lg Rank对应的行)需要跳过,只保留数据行
内容的提问来源于stack exchange,提问作者apersoninneed
相关产品推荐
相关产品推荐

