Pandas滚动窗口:序列首尾保留窗口宽度的实现问询
Absolutely, you can get this exact behavior in pandas! The key is to adjust how the rolling window handles incomplete boundary windows so you don't get NaNs at the start/end, while keeping full-width windows for the middle of your sequence.
Here are two straightforward approaches:
1. Use min_periods=1 with rolling()
This is the simplest method. By setting min_periods=1, you tell pandas to compute the statistic even if the window doesn't have the full number of observations—instead, it uses whatever data is available in the window. For the middle of your sequence, the window will still be full-width, and at the ends, it will use the maximum possible subset of data (growing from 1 to your window size at the start, shrinking from full size to 1 at the end).
Example code for a centered window with width 3:
import pandas as pd import numpy as np # Sample data data = pd.Series([1, 2, 3, 4, 5, 6, 7]) # Rolling window with center=True and min_periods=1 rolling_centered = data.rolling(window=3, center=True, min_periods=1).mean() print(rolling_centered)
Output:
0 1.0 1 2.0 2 3.0 3 4.0 4 5.0 5 6.0 6 7.0 dtype: float64
Without min_periods=1, the first and last values would be NaN. Now, the middle values (indices 1-5) use the full 3-element window, while the first and last use only their own single value.
2. Custom Boundary Handling (For More Control)
If you want even more control over how the boundaries are calculated (e.g., using a fixed partial window size instead of growing/shrinking), you can manually compute the boundary statistics and combine them with the full-window results.
For example, let's say you want the first 2 elements (for a window size of 3) to use a 2-element window, and the last 2 to use a 2-element window:
window_size = 3 half_window = (window_size - 1) // 2 # Full rolling window (will have NaNs at boundaries) full_rolling = data.rolling(window=window_size, center=True).mean() # Compute start boundary stats (first half_window elements) start_stats = data.iloc[:half_window + 1].expanding().mean() # Compute end boundary stats (last half_window elements) end_stats = data.iloc[-half_window - 1:].iloc[::-1].expanding().mean().iloc[::-1] # Combine everything custom_rolling = full_rolling.copy() custom_rolling.iloc[:half_window] = start_stats.iloc[:half_window] custom_rolling.iloc[-half_window:] = end_stats.iloc[-half_window:] print(custom_rolling)
This gives you precise control over how the edge cases are handled, while keeping the full window width for the main part of the sequence.
Key Notes
- When using
min_periods=1, make sure this aligns with your statistical needs—using partial windows at the ends will change the statistic values compared to full windows, but it eliminates NaNs. - For centered windows, the number of boundary elements on each end is
(window_size - 1) // 2(e.g., window size 5 has 2 boundary elements on each side).
内容的提问来源于stack exchange,提问作者ascripter

