You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于递增增长窗口遍历DataFrame Out并应用规则的技术问询

Expanding Window Behaviour Classification for DataFrame Out

Got it, let's walk through how to implement this expanding window logic for your DataFrame Out—I'll break this down into clear, actionable steps with code examples, and call out any ambiguous bits I notice along the way.

First, Let's Align on the Window Definition

You mentioned an expanding window where the tail stays fixed, and the head expands backward (adding earlier rows each iteration) to make the window larger each time. So for example, if we're focusing on row 5 as the fixed tail, we'd start with just row 5, then rows 4-5, then 3-5, and so on until one of your behaviour rules is satisfied.

The Behaviour Rules (Clarified)

From your description, here's the order of checks we need to run on each expanding window:

  1. If the mean of Speed_out in the window is ≤ 0.152682 → mark the tail row's Behaviour as Lying
  2. If the first condition fails, check two things:
    • Movement_Out (I assume this is the value from the fixed tail row? Or a window-wide value? I'll note how to adjust this later) equals "left" and that value is ≤20.8
    • Plus, the mean of Speed_Out in the window is ≥0.200921
    • If both are true → mark Behaviour as Grazing
  3. If neither condition is met, keep expanding the window by adding one more earlier row and repeat.

Code Implementation (Using Pandas)

Here's a practical, runnable example. I'll use a sample DataFrame to demonstrate, but you can swap it out with your actual Out data:

import pandas as pd

# Sample DataFrame (replace this with your real `Out` data)
Out = pd.DataFrame({
    'Speed_out': [0.1, 0.13, 0.22, 0.19, 0.26],
    'Movement_Out': ["left", "right", "left", "left", "right"]
})

# Initialize the Behaviour column with missing values
Out['Behaviour'] = pd.NA

# Loop over each row to use as the fixed window tail
for tail_idx in range(len(Out)):
    # Expand the window backward: start with just the tail, then add earlier rows one by one
    for start_idx in range(tail_idx, -1, -1):
        # Slice the current expanding window
        current_window = Out.loc[start_idx:tail_idx, :]
        speed_avg = current_window['Speed_out'].mean()
        
        # Check first condition: Lying
        if speed_avg <= 0.152682:
            Out.loc[tail_idx, 'Behaviour'] = 'Lying'
            break  # Stop expanding once we hit a matching condition
        
        # Check second condition: Grazing
        # Important note: If `Movement_Out` is a string like "left", comparing to 20.8 doesn't make sense
        # I'm assuming this might be a typo—maybe you meant a numeric column? Or that `Movement_Out` is numeric and "left" was a mistake?
        # For now, I'll follow your logic, but adjust this part to match your actual data structure
        tail_movement = Out.loc[tail_idx, 'Movement_Out']
        if tail_movement == "left" and float(tail_movement) <= 20.8 and speed_avg >= 0.200921:
            Out.loc[tail_idx, 'Behaviour'] = 'Grazing'
            break  # Stop expanding once condition is met

# Print the final result
print(Out)

Key Adjustments You Might Need:

  • Movement_Out Ambiguity: If Movement_Out is a categorical string column (like "left"/"right"), the <=20.8 check doesn't make sense. You probably meant to reference a different numeric column here, or maybe Movement_Out is numeric and "left" was a typo. Fix this line to match your actual data.
  • Window-Based Movement_Out Check: If you need to check the Movement_Out value across the entire window (e.g., all values are "left", or the mean is ≤20.8), replace tail_movement with current_window['Movement_Out'].some_aggregation() (like .mean() or .unique()).
  • Performance for Large DataFrames: The nested loop works for small to medium DataFrames, but for large datasets, you can optimize using pandas' built-in expanding() method with a custom apply function to avoid looping row-by-row.
  • Default Behaviour: If a window never meets either condition, the Behaviour column will stay as pd.NA. Add an else clause at the end of the inner loop to set a default (like 'Unknown') if needed.

Edge Cases to Keep in Mind

  • Ensure Speed_out is a numeric column—if it's stored as strings, convert it first with Out['Speed_out'] = pd.to_numeric(Out['Speed_out'], errors='coerce').
  • If your DataFrame has missing values, decide how to handle them (e.g., use mean(skipna=True) which is pandas' default, or drop NaNs first).

内容的提问来源于stack exchange,提问作者PharmR

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:47:50