如何通过循环基于DataFrame特定列值识别实例并截取'Shot'对应行及前4行生成序列
Got it, let's tackle this problem step by step using pandas—perfect for your 1000-row DataFrame. Here's how to efficiently get the rows for each 'Shot' event plus the 4 preceding rows:
Step 1: Locate all 'Shot' event indices
First, grab the index positions of every row where eventName equals 'Shot':
import pandas as pd import numpy as np # Assume your DataFrame is named df shot_indices = df[df['eventName'] == 'Shot'].index
Step 2: Gather all relevant indices
We need to include each Shot's index, plus the 4 indices before it (or start from row 0 if there aren't 4 prior rows). Using numpy keeps this clean and automatically handles overlapping ranges to avoid duplicate rows:
# Generate an array of all indices we need to keep all_relevant_indices = np.unique( np.concatenate([np.arange(max(0, idx - 4), idx + 1) for idx in shot_indices]) )
The max(0, idx -4) prevents us from trying to access negative indices for shots that occur in the first 4 rows. np.unique removes duplicates that might come from overlapping ranges (like two consecutive shots sharing some preceding rows).
Step 3: Slice the DataFrame for your final result
Now just extract the rows using the collected indices:
result_df = df.loc[all_relevant_indices]
Example with your sample data
In your provided sample, the 'Shot' is at index 4. The code will select indices 0-4, returning all 5 rows of your sample—exactly what you need.
Alternative loop-based approach (for clarity)
If you prefer a more explicit, beginner-friendly loop, this works too:
selected_rows = [] for idx in shot_indices: # Ensure we don't go below row 0 start_idx = max(0, idx - 4) # Append the slice from start_idx to idx (inclusive) selected_rows.append(df.loc[start_idx:idx]) # Combine slices and remove duplicate rows result_df = pd.concat(selected_rows).drop_duplicates()
Both methods will give you the sequence of rows you need for further processing.
内容的提问来源于stack exchange,提问作者Stuart Macfarlane

