You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于Pandas实现从距今90天前到30天前的Rolling mean?

Got it, let's tackle this problem. You want to adjust your rolling mean from the last 90 days up to today, to the window between 90 days ago and 30 days ago—so each value corresponds to the average quantity in that 60-day gap (90-30=60 days) prior to the current date.

Here are two straightforward ways to implement this in pandas:

Method 1: Use Rolling + Shift (Optimized for Speed)

The core idea is to leverage pandas' built-in rolling operations combined with a shift to adjust the window position:

  1. Calculate a 60-day rolling mean (since the window between 90d and 30d ago is exactly 60 days long)
  2. Shift the result forward by 30 days, so that each date's value now reflects the average from 90d to 30d before it.

Code example:

# First ensure your Date column is datetime type (critical for time-based rolling)
df['Date'] = pd.to_datetime(df['Date'])

# Compute the 60-day rolling mean, then shift 30 days to align with the desired window
df['rolling_mean_90_to_30'] = df.rolling(window='60d', on='Date')['quantity'].mean().shift(periods='30d')

Method 2: Explicit Window Filtering (For Full Control)

If you prefer a more readable, explicit approach (great for debugging or adding custom logic), you can use apply to filter the exact date range for each row:

df['Date'] = pd.to_datetime(df['Date'])

def calculate_90_30_mean(row):
    # Define the start and end of the target window for the current row's date
    window_start = row['Date'] - pd.Timedelta(days=90)
    window_end = row['Date'] - pd.Timedelta(days=30)
    # Filter rows within the window and compute the mean
    window_data = df[(df['Date'] >= window_start) & (df['Date'] <= window_end)]['quantity']
    return window_data.mean()

# Apply the custom function to each row
df['rolling_mean_90_to_30'] = df.apply(calculate_90_30_mean, axis=1)

Key Notes:

  • Performance: Method 1 is far faster for large datasets, as it uses pandas' optimized vectorized operations. Method 2 is more intuitive but can be slow with big data due to row-wise processing.
  • Missing Values: Both methods return NaN for dates where there isn't enough historical data (e.g., dates less than 90 days after your dataset's earliest date), matching the behavior of your original rolling code.
  • Window Inclusivity: If you need to tweak whether the start/end dates are included in the calculation, add the closed parameter to the rolling function (e.g., closed='left' to exclude the window's end date).

内容的提问来源于stack exchange,提问作者william007

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:48:06