You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中对整个DataFrame应用滚动自定义函数?

Great question! The default df.rolling(N).apply(f) processes data column-wise, which isn't ideal when you need to operate on the entire rolling window of the DataFrame (all columns together). Here are a couple of straightforward ways to achieve this:

Method 1: Use NumPy's Sliding Window View (Efficient for Large Data)

This approach leverages NumPy's sliding_window_view to extract all rolling windows as 2D sub-arrays, then applies your custom function to each window. It’s often faster than pure pandas methods for big datasets.

Step-by-Step Example:

  1. Import required libraries:
import pandas as pd
import numpy as np
  1. Create a sample DataFrame:
df = pd.DataFrame({
    'A': [1, 2, 3, 4, 5],
    'B': [6, 7, 8, 9, 10],
    'C': [11, 12, 13, 14, 15]
})
  1. Define your custom function that takes a full window DataFrame:
def process_full_window(window_df):
    # Example: Calculate the sum of all column means in the window
    col_means = window_df.mean()
    return col_means.sum()
  1. Generate sliding windows and apply the function:
window_size = 2

# Get sliding window view (shape: (number_of_windows, window_size, number_of_columns))
windows = np.lib.stride_tricks.sliding_window_view(df.values, window_shape=(window_size, df.shape[1]))

# Apply the function to each window (convert each window array back to DataFrame first)
results = [process_full_window(pd.DataFrame(window)) for window in windows]

# Prepend NaNs for the initial rows that don't have a full window
final_results = pd.Series(
    [np.nan] * (window_size - 1) + results,
    index=df.index
)

print(final_results)

Output:

0     NaN
1    21.0
2    27.0
3    33.0
4    39.0
dtype: float64

Method 2: Pandas Rolling Apply with Reshaping (API-Centric)

If you prefer to stick with pandas' rolling API, you can use raw=True to get window data as NumPy arrays, then reshape them back into a DataFrame for your function. Note that this will return the same result across all columns (since we’re processing the full window once), so you might want to drop duplicate columns afterward.

Example:

def process_full_window(window_df):
    # Same example function as above
    col_means = window_df.mean()
    return col_means.sum()

window_size = 2

# Use rolling.apply with raw=True to get window arrays, then reshape
result_df = df.rolling(window_size).apply(
    lambda win_array: process_full_window(pd.DataFrame(win_array.reshape(-1, df.shape[1]))),
    raw=True
)

# Since all columns have the same result, you can keep just one if needed
final_result = result_df.iloc[:, 0]

print(final_result)

This will give the same output as the NumPy method.

Key Notes:

  • If your function returns a Series instead of a scalar, the NumPy method can be adjusted to collect those Series into a final DataFrame.
  • For very large datasets, the NumPy sliding window method is usually more efficient because it avoids some of pandas' overhead.

内容的提问来源于stack exchange,提问作者Alex

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:19:45