You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python解析XML字符串提取投篮数据生成DataFrame方法

问题根因

代码返回None由两个问题导致:

  • XML根节点声明了默认命名空间xmlns="http://feed.elasticstats.com/schema/basketball/pbp-v7.0.xsd",xml.etree.ElementTree查找带默认命名空间的节点时,必须显式传入命名空间映射,给标签名拼接命名空间前缀,否则无法匹配到任何节点。
  • 原查找逻辑只取第一个匹配的event节点,无法拿到全量事件数据。
可直接运行的实现代码
import xml.etree.ElementTree as et
import pandas as pd

# 注册XML默认命名空间
NS = {"pbp": "http://feed.elasticstats.com/schema/basketball/pbp-v7.0.xsd"}
root = et.fromstring(xml)

shot_data = []
# 递归查找所有节次下的事件节点,兼容多节、加时赛场景
for event in root.findall(".//pbp:quarter/pbp:events/pbp:event", NS):
    # 过滤无坐标、无投篮统计的非投篮事件(换人、犯规、跳球等)
    loc = event.find("pbp:location", NS)
    field_goal = event.find("pbp:statistics/pbp:fieldgoal", NS)
    shooter = event.find("pbp:statistics/pbp:fieldgoal/pbp:player", NS)
    if not all([loc, field_goal, shooter]):
        continue

    shot_data.append({
        "球员姓名": shooter.get("full_name"),
        "投篮位置": f"({loc.get('coord_x')}, {loc.get('coord_y')})",
        "coord_x": int(loc.get("coord_x")),
        "coord_y": int(loc.get("coord_y")),
        "时间": event.get("clock"),
        "投篮类型": field_goal.get("shot_type"),
        "事件类型": event.get("event_type"),
        "是否命中": field_goal.get("made") == "true",
        "出手区域": loc.get("action_area"),
        "投篮距离(英尺)": float(field_goal.get("shot_distance"))
    })

# 生成目标DataFrame,以投篮球员为行索引
shot_df = pd.DataFrame(shot_data).set_index("球员姓名")
注意事项
  • 路径中使用.//前缀做递归查找,不需要手动遍历每一个quarter节点,自动覆盖常规4节+加时赛的所有事件
  • 非空判断会自动跳过无坐标的事件,比如样例中第一个lineupchange换人事件会被直接过滤,不会污染投篮数据集
  • 数值类字段(坐标、投篮距离)直接做了类型转换,后续做投篮热图绘制、统计分析不需要额外转格式
  • 如果只需要要求的三个列(投篮位置、时间、投篮类型),删掉字典里多余的键值对即可。

内容的提问来源于stack exchange,提问作者Nick

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 16:01:05