如何用nba_api分析球员单场抢断泊松分布?代码报错求助
NBA球员单场抢断泊松分布验证问题修复
问题背景
目标验证NBA球员单场抢断数据是否符合泊松分布:按赛季场均抢断的小数数值分组(例如场均1.2次),统计对应赛季中单场0次、1次等抢断的场次分布,通过方差和直方图对比λ=1.2的泊松分布。但编写的代码出现两种问题:一是长时间无响应后触发连接/运行时错误,二是抛出ValueError: No objects to concatenate。
错误原因
- 精确匹配无数据:代码中用
career_stats_df['STL'] == 1.2精确筛选场均抢断,而场均抢断是总抢断数除以出场场次的结果,几乎没有球员刚好等于1.2,导致player_seasons和player_game_logs为空,拼接时触发错误。 - API限流与请求过载:遍历所有NBA球员并发起大量API请求,容易触发官方接口的限流机制,导致超时或连接失败。
- 未处理异常请求:部分球员可能无有效生涯数据或比赛日志,直接调用
get_data_frames()会引发错误,中断程序。
修复后的代码
from nba_api.stats.static import players from nba_api.stats.endpoints import playercareerstats, playergamelog import pandas as pd import time import random # 筛选过去10个赛季(2013-14到2023-24) target_seasons = [f"{year}-{str(year+1)[-2:]}" for year in range(2013, 2024)] # 获取近年活跃球员(现役或近10年有参赛记录) nba_players = [p for p in players.get_players() if p['is_active'] or int(p['last_played']) >= 2013] player_seasons = [] player_game_logs = [] for idx, player in enumerate(nba_players): player_id = player['id'] # 每处理10个球员添加随机延迟,规避限流 if idx % 10 == 0 and idx != 0: time.sleep(random.uniform(1, 3)) try: # 获取球员生涯数据 career_stats = playercareerstats.PlayerCareerStats(player_id=player_id) career_stats_df = career_stats.get_data_frames()[0] # 筛选场均抢断在1.15-1.25之间的目标赛季 filtered_seasons = career_stats_df[ (career_stats_df['STL'].between(1.15, 1.25)) & (career_stats_df['SEASON_ID'].isin(target_seasons)) ] if not filtered_seasons.empty: player_seasons.append(filtered_seasons) for season in filtered_seasons['SEASON_ID']: # 请求单赛季日志前添加小延迟 time.sleep(random.uniform(0.5, 1.5)) game_log = playergamelog.PlayerGameLog(player_id=player_id, season=season) game_log_df = game_log.get_data_frames()[0] # 补充球员和赛季标识,方便后续分析 game_log_df['PLAYER_ID'] = player_id game_log_df['SEASON_ID'] = season player_game_logs.append(game_log_df) except Exception as e: print(f"处理球员{player['full_name']}时出错: {e}") continue # 拼接并输出结果 if player_seasons: result_seasons_df = pd.concat(player_seasons, ignore_index=True) print(f"找到符合条件的赛季数: {len(result_seasons_df)}") else: print("未找到符合条件的赛季数据") if player_game_logs: result_game_logs_df = pd.concat(player_game_logs, ignore_index=True) # 统计单场抢断各数值的场次分布 steal_distribution = result_game_logs_df['STL'].value_counts().sort_index() print("\n单场抢断场次分布:") print(steal_distribution) # 统计0次抢断总场次 total_zero_steals = len(result_game_logs_df[result_game_logs_df['STL'] == 0]) print(f"\n0次抢断总场次: {total_zero_steals}") else: print("未找到符合条件的比赛日志数据")
关键优化点
- 范围匹配代替精确匹配:用
between(1.15, 1.25)筛选场均抢断接近1.2的赛季,避免浮点精度导致无数据。 - 限流控制:添加随机延迟,降低API请求频率,规避官方限流机制。
- 缩小球员范围:只处理近年活跃球员,减少无效请求数量。
- 异常捕获:处理请求过程中的异常,避免单个球员出错导致程序中断。
- 数据增强:给比赛日志补充球员ID和赛季ID,方便后续分组分析。
内容的提问来源于stack exchange,提问作者SpeedyClaxton
相关产品推荐
相关产品推荐

