You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用nba_api分析球员单场抢断泊松分布?代码报错求助

NBA球员单场抢断泊松分布验证问题修复

问题背景

目标验证NBA球员单场抢断数据是否符合泊松分布:按赛季场均抢断的小数数值分组(例如场均1.2次),统计对应赛季中单场0次、1次等抢断的场次分布,通过方差和直方图对比λ=1.2的泊松分布。但编写的代码出现两种问题:一是长时间无响应后触发连接/运行时错误,二是抛出ValueError: No objects to concatenate。

错误原因

  1. 精确匹配无数据:代码中用career_stats_df['STL'] == 1.2精确筛选场均抢断,而场均抢断是总抢断数除以出场场次的结果,几乎没有球员刚好等于1.2,导致player_seasons和player_game_logs为空,拼接时触发错误。
  2. API限流与请求过载:遍历所有NBA球员并发起大量API请求,容易触发官方接口的限流机制,导致超时或连接失败。
  3. 未处理异常请求:部分球员可能无有效生涯数据或比赛日志,直接调用get_data_frames()会引发错误,中断程序。

修复后的代码

from nba_api.stats.static import players
from nba_api.stats.endpoints import playercareerstats, playergamelog
import pandas as pd
import time
import random

# 筛选过去10个赛季(2013-14到2023-24)
target_seasons = [f"{year}-{str(year+1)[-2:]}" for year in range(2013, 2024)]

# 获取近年活跃球员(现役或近10年有参赛记录)
nba_players = [p for p in players.get_players() if p['is_active'] or int(p['last_played']) >= 2013]

player_seasons = []
player_game_logs = []

for idx, player in enumerate(nba_players):
    player_id = player['id']
    
    # 每处理10个球员添加随机延迟,规避限流
    if idx % 10 == 0 and idx != 0:
        time.sleep(random.uniform(1, 3))
    
    try:
        # 获取球员生涯数据
        career_stats = playercareerstats.PlayerCareerStats(player_id=player_id)
        career_stats_df = career_stats.get_data_frames()[0]
        
        # 筛选场均抢断在1.15-1.25之间的目标赛季
        filtered_seasons = career_stats_df[
            (career_stats_df['STL'].between(1.15, 1.25)) &
            (career_stats_df['SEASON_ID'].isin(target_seasons))
        ]
        
        if not filtered_seasons.empty:
            player_seasons.append(filtered_seasons)
            
            for season in filtered_seasons['SEASON_ID']:
                # 请求单赛季日志前添加小延迟
                time.sleep(random.uniform(0.5, 1.5))
                game_log = playergamelog.PlayerGameLog(player_id=player_id, season=season)
                game_log_df = game_log.get_data_frames()[0]
                # 补充球员和赛季标识,方便后续分析
                game_log_df['PLAYER_ID'] = player_id
                game_log_df['SEASON_ID'] = season
                player_game_logs.append(game_log_df)
                
    except Exception as e:
        print(f"处理球员{player['full_name']}时出错: {e}")
        continue

# 拼接并输出结果
if player_seasons:
    result_seasons_df = pd.concat(player_seasons, ignore_index=True)
    print(f"找到符合条件的赛季数: {len(result_seasons_df)}")
else:
    print("未找到符合条件的赛季数据")

if player_game_logs:
    result_game_logs_df = pd.concat(player_game_logs, ignore_index=True)
    # 统计单场抢断各数值的场次分布
    steal_distribution = result_game_logs_df['STL'].value_counts().sort_index()
    print("\n单场抢断场次分布:")
    print(steal_distribution)
    
    # 统计0次抢断总场次
    total_zero_steals = len(result_game_logs_df[result_game_logs_df['STL'] == 0])
    print(f"\n0次抢断总场次: {total_zero_steals}")
else:
    print("未找到符合条件的比赛日志数据")

关键优化点

  • 范围匹配代替精确匹配:用between(1.15, 1.25)筛选场均抢断接近1.2的赛季,避免浮点精度导致无数据。
  • 限流控制:添加随机延迟,降低API请求频率,规避官方限流机制。
  • 缩小球员范围:只处理近年活跃球员,减少无效请求数量。
  • 异常捕获:处理请求过程中的异常,避免单个球员出错导致程序中断。
  • 数据增强:给比赛日志补充球员ID和赛季ID,方便后续分组分析。

内容的提问来源于stack exchange,提问作者SpeedyClaxton

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 09:12:49