将计算特定移动平均的R循环代码转换为Python实现的技术请求
Got it, let's translate your R iteration logic to Python—we'll cover both a clean, efficient pandas approach (the best practice for tabular data) and a direct loop replication that matches your original code step-by-step.
Core Logic Breakdown
First, let's recap what your R code is doing to make sure we don't miss anything:
- You're grouping rows by the
Playercolumn - For each player's sequence of observations:
- The
mintotcolumn is 0 for the player's first row, then for every subsequent row it's the sum ofMINvalues from the player's first row up to the previous row - For columns 16 through 78 (1-based in R), you calculate the cumulative sum of each column up to the previous row, divide by
mintot, and store the result in columns shifted 65 positions to the right
- The
- You print progress every time you finish processing 10 player groups
Efficient Pandas Implementation
Pandas is built for this kind of grouped, cumulative operation—vectorized methods will run way faster than nested loops, especially with large datasets. Here's how to do it:
Step 1: Calculate mintot
We'll use grouped cumulative sums and shift the result to get the sum up to the previous row, then fill the first row's missing value with 0:
import pandas as pd # Assuming your data is loaded into a pandas DataFrame called df df['mintot'] = df.groupby('Player')['MIN'].cumsum().shift(1).fillna(0)
Step 2: Compute the Moving Averages
First, map R's 1-based column indices to Python's 0-based system:
- R's columns 16–78 → Python's indices 15–77
- R's target columns (j+65) → Python's indices 80–142 (since 15+65=80, 77+65=142)
Then loop through each source-target column pair to compute the cumulative average:
# Define source and target columns source_cols = df.columns[15:78] # R's 16-78 target_cols = df.columns[80:143] # R's j+65 (16+65=81 to 78+65=143, which is 80-142 in 0-based) for src_col, tgt_col in zip(source_cols, target_cols): # Get cumulative sum per player, shifted to exclude the current row cum_sum = df.groupby('Player')[src_col].cumsum().shift(1).fillna(0) # Divide by mintot, handle division by 0 (set to NaN or 0 to match R's behavior) df[tgt_col] = cum_sum / df['mintot'].replace(0, float('nan')) # Use this line instead if you want 0 instead of NaN for division by 0: # df[tgt_col] = cum_sum / df['mintot'].replace(0, float('nan')).fillna(0)
Step 3: Progress Tracking
If you want to print progress every 10 players, you can iterate through the groups directly:
player_groups = df.groupby('Player') total_players = len(player_groups) for idx, (player, _) in enumerate(player_groups, 1): if idx % 10 == 0: print(f"Processed {idx} of {total_players} players")
Direct Loop Replication (Matching R's Exact Flow)
If you want a line-for-line match to your R code (useful for debugging or understanding the iteration), here's the nested loop version (note Python uses 0-based indexing, so we adjust row numbers accordingly):
k = 0 # Start at 0 instead of 1 (Python is 0-based) n_rows = len(df) while k < n_rows: i = 0 current_player = df.iloc[k]['Player'] # Loop through all rows for the current player while k + i < n_rows and df.iloc[k + i]['Player'] == current_player: if i == 0: # First row of the player: mintot = 0 df.loc[k + i, 'mintot'] = 0 else: # Sum MIN from k to k+i-1 (Python slicing is left-closed, right-open) df.loc[k + i, 'mintot'] = df.iloc[k:k+i]['MIN'].sum() # Process columns 16-78 (R) → 15-77 (Python) for j in range(15, 78): sum_val = df.iloc[k:k+i, j].sum() mintot_val = df.loc[k + i, 'mintot'] if mintot_val == 0: # Handle division by 0 (matches R's behavior of 0/0 → NaN) df.iloc[k + i, j + 65] = float('nan') else: df.iloc[k + i, j + 65] = sum_val / mintot_val i += 1 k += i if k % 10 == 0: print(k)
Notes on Indexing & Edge Cases
- Division by 0: The first row of each player has
mintot = 0, so the average calculation will result inNaN(or 0 if you use the fill option). This matches your original R code, wheresum(df$MIN[k:(k+i-1)])would be 0 for the first row, leading to 0/0. - Performance: The pandas vectorized method is vastly faster than nested loops for large datasets—use the loop version only if you need to replicate the exact iteration flow for debugging.
内容的提问来源于stack exchange,提问作者Riccardo Tornaghi

