基于篮球DataFrame按PlayId生成近距球员位置与速度列的技术问询
Solution to Add Closest Players' Pos and Speed for Ball Handlers (Multi PlayId Support)
Hey there, let's tackle this problem step by step. You already have the DistanceToBall column calculated and filtered down to ball handlers in newdf. Now we need to, for each PlayId, grab the 4 closest non-ball-handler players' Pos and Speed values and append them as columns to the ball handler's row—here's a robust approach that works seamlessly with multiple plays:
Step 1: Core Logic Overview
We'll use groupby('PlayId') to handle each play independently. For every play group:
- Isolate the single ball handler row (where
Ball == 1) - Sort remaining players by
DistanceToBall(ascending, so closest first) - Extract the top 4 entries'
PosandSpeed, reshape them into columns, and merge with the ball handler's data
Step 2: Full Code Implementation
import pandas as pd import numpy as np # Sample data (expanded to test multi-PlayId scenario) df = pd.DataFrame({ 'PlayId': [1,1,1,1,1,2,2,2,2,2], 'Player': ['A','B','C','D','E','F','G','H','I','J'], 'Ball': [0,0,1,0,0,0,1,0,0,0], 'Pos': [1,4,10,15,20,5,12,8,18,25], 'Speed': [1,2,3,4,5,6,7,8,9,10] }) # Your existing DistanceToBall calculation df['DistanceToBall'] = np.abs(df.Pos - df.Pos[df.groupby('PlayId')['Ball'].transform('idxmax')].reset_index(drop=True)) # Function to process each play group def process_single_play(group): # Grab the ball handler's row ball_handler = group[group['Ball'] == 1].copy() # Get non-ball-handler players, sort by closeness other_players = group[group['Ball'] != 1].sort_values('DistanceToBall', ascending=True) # Extract top 4 Pos/Speed, pad with NaN if fewer than 4 players exist closest_pos = other_players['Pos'].head(4).tolist() closest_speed = other_players['Speed'].head(4).tolist() closest_pos += [np.nan] * (4 - len(closest_pos)) closest_speed += [np.nan] * (4 - len(closest_speed)) # Attach new columns to the ball handler row ball_handler[['closest1', 'closest2', 'closest3', 'closest4']] = closest_pos ball_handler[['speed1', 'speed2', 'speed3', 'speed4']] = closest_speed return ball_handler # Apply to all plays and combine results newdf = df.groupby('PlayId').apply(process_single_play).reset_index(drop=True) # View the final output print(newdf)
Step 3: Key Details Explained
- Multi-Play Safety: Grouping by
PlayIdensures we never mix data across different plays, which is critical for scalability. - Sorting & Padding: Sorting by
DistanceToBallguarantees we get the closest players first. We pad withNaNif a play has fewer than 4 non-ball-handler players to keep the DataFrame structure consistent. - Expected Output Match: For your original PlayId=1 sample, the ball handler row will have
closest1-4as15,4,1,20andspeed1-4as4,2,1,5—exactly matching your desired result.
内容的提问来源于stack exchange,提问作者yankeefan11
相关产品推荐
相关产品推荐

