基于递增增长窗口遍历DataFrame Out并应用规则的技术问询
Out Got it, let's walk through how to implement this expanding window logic for your DataFrame Out—I'll break this down into clear, actionable steps with code examples, and call out any ambiguous bits I notice along the way.
First, Let's Align on the Window Definition
You mentioned an expanding window where the tail stays fixed, and the head expands backward (adding earlier rows each iteration) to make the window larger each time. So for example, if we're focusing on row 5 as the fixed tail, we'd start with just row 5, then rows 4-5, then 3-5, and so on until one of your behaviour rules is satisfied.
The Behaviour Rules (Clarified)
From your description, here's the order of checks we need to run on each expanding window:
- If the mean of
Speed_outin the window is ≤ 0.152682 → mark the tail row'sBehaviourasLying - If the first condition fails, check two things:
Movement_Out(I assume this is the value from the fixed tail row? Or a window-wide value? I'll note how to adjust this later) equals "left" and that value is ≤20.8- Plus, the mean of
Speed_Outin the window is ≥0.200921 - If both are true → mark
BehaviourasGrazing
- If neither condition is met, keep expanding the window by adding one more earlier row and repeat.
Code Implementation (Using Pandas)
Here's a practical, runnable example. I'll use a sample DataFrame to demonstrate, but you can swap it out with your actual Out data:
import pandas as pd # Sample DataFrame (replace this with your real `Out` data) Out = pd.DataFrame({ 'Speed_out': [0.1, 0.13, 0.22, 0.19, 0.26], 'Movement_Out': ["left", "right", "left", "left", "right"] }) # Initialize the Behaviour column with missing values Out['Behaviour'] = pd.NA # Loop over each row to use as the fixed window tail for tail_idx in range(len(Out)): # Expand the window backward: start with just the tail, then add earlier rows one by one for start_idx in range(tail_idx, -1, -1): # Slice the current expanding window current_window = Out.loc[start_idx:tail_idx, :] speed_avg = current_window['Speed_out'].mean() # Check first condition: Lying if speed_avg <= 0.152682: Out.loc[tail_idx, 'Behaviour'] = 'Lying' break # Stop expanding once we hit a matching condition # Check second condition: Grazing # Important note: If `Movement_Out` is a string like "left", comparing to 20.8 doesn't make sense # I'm assuming this might be a typo—maybe you meant a numeric column? Or that `Movement_Out` is numeric and "left" was a mistake? # For now, I'll follow your logic, but adjust this part to match your actual data structure tail_movement = Out.loc[tail_idx, 'Movement_Out'] if tail_movement == "left" and float(tail_movement) <= 20.8 and speed_avg >= 0.200921: Out.loc[tail_idx, 'Behaviour'] = 'Grazing' break # Stop expanding once condition is met # Print the final result print(Out)
Key Adjustments You Might Need:
Movement_OutAmbiguity: IfMovement_Outis a categorical string column (like "left"/"right"), the<=20.8check doesn't make sense. You probably meant to reference a different numeric column here, or maybeMovement_Outis numeric and "left" was a typo. Fix this line to match your actual data.- Window-Based
Movement_OutCheck: If you need to check theMovement_Outvalue across the entire window (e.g., all values are "left", or the mean is ≤20.8), replacetail_movementwithcurrent_window['Movement_Out'].some_aggregation()(like.mean()or.unique()). - Performance for Large DataFrames: The nested loop works for small to medium DataFrames, but for large datasets, you can optimize using pandas' built-in
expanding()method with a customapplyfunction to avoid looping row-by-row. - Default Behaviour: If a window never meets either condition, the
Behaviourcolumn will stay aspd.NA. Add an else clause at the end of the inner loop to set a default (like'Unknown') if needed.
Edge Cases to Keep in Mind
- Ensure
Speed_outis a numeric column—if it's stored as strings, convert it first withOut['Speed_out'] = pd.to_numeric(Out['Speed_out'], errors='coerce'). - If your DataFrame has missing values, decide how to handle them (e.g., use
mean(skipna=True)which is pandas' default, or drop NaNs first).
内容的提问来源于stack exchange,提问作者PharmR

