You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python机器学习结合Pandas DataFrame预测NFL跑传类型?是否用决策树?

NFL比赛回合类型(跑球/传球)预测实现方案

一、是否适合用scikit-learn决策树?

完全可以用,而且是很合适的选择。决策树的核心优势是可解释性极强——训练完成后,你能直观看到哪些特征(比如4档、需要推进的码数多)会让模型倾向于选择传球,刚好匹配你想理解决策逻辑的需求。如果追求更高的预测精度,也可以试试随机森林、XGBoost这类集成模型,它们的泛化能力会比单棵决策树更好。

二、完整实现步骤

1. 数据清洗与预处理

首先要过滤无效样本,提取需要的特征和目标变量:

import nfl_data_py as nfl
import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier
from sklearn.preprocessing import LabelEncoder
from sklearn.metrics import accuracy_score, classification_report

# 导入2022赛季数据
pbp = nfl.import_pbp_data([2022], downcast=True, cache=False, alt_path=None)

# 筛选有效样本:仅保留play_type为run/pass的回合,且所有特征无缺失值
filtered_pbp = pbp[
    pbp['play_type'].isin(['run', 'pass']) &
    pbp[['score_differential', 'yardline_100', 'ydstogo', 'down', 'half_seconds_remaining']].notna().all(axis=1)
].copy()

# 对目标变量编码:将run/pass转为模型可识别的数值(run=0,pass=1)
le = LabelEncoder()
filtered_pbp['play_type_encoded'] = le.fit_transform(filtered_pbp['play_type'])

# 定义特征矩阵X和目标向量y
X = filtered_pbp[['score_differential', 'yardline_100', 'ydstogo', 'down', 'half_seconds_remaining']]
y = filtered_pbp['play_type_encoded']

2. 划分训练集与测试集

按8:2的比例拆分数据,用于模型训练和效果验证:

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

3. 训练决策树模型

初始化模型并训练,可通过max_depth参数限制树的深度,避免过拟合:

dt_model = DecisionTreeClassifier(max_depth=5, random_state=42)
dt_model.fit(X_train, y_train)

4. 模型效果评估

在测试集上验证模型的预测能力:

# 生成测试集预测结果
y_pred = dt_model.predict(X_test)

# 输出准确率和分类报告
print(f"模型准确率: {accuracy_score(y_test, y_pred):.2f}")
print("\n分类报告:")
print(classification_report(y_test, y_pred, target_names=le.classes_))

5. 自定义输入预测

针对你提到的场景(分差-4、码线25、4档、需推进16码、半场剩余300秒),直接构造特征数据即可得到预测结果:

# 构造自定义输入,顺序与特征列表一致
custom_input = pd.DataFrame([[-4, 25, 16, 4, 300]], columns=X.columns)

# 预测并解码回原始类别
predicted_code = dt_model.predict(custom_input)
predicted_play_type = le.inverse_transform(predicted_code)[0]

print(f"预测的回合类型: {predicted_play_type}")

三、优化建议

  • 特征工程:把down(档数)转为分类特征(用OneHotEncoder),因为档数是离散类别而非连续数值;也可以对half_seconds_remaining做分箱处理(比如<60秒、60-300秒、>300秒),让模型更好捕捉时间压力的影响。
  • 参数调优:用GridSearchCV或RandomizedSearchCV优化决策树的max_depth、min_samples_split等参数,进一步提升模型性能。
  • 模型对比:尝试随机森林、XGBoost这类集成模型,它们的预测稳定性和精度通常优于单棵决策树。

内容的提问来源于stack exchange,提问作者earningjoker430

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 15:10:39