Rolling Apply为何卡顿?大型DataFrame滚动加权均值实现求助
rolling.apply Is Slow Hey there! Let's tackle your two main questions: why rolling.apply is so sluggish, and how to efficiently compute those triangular weighted averages for your 100M-row DataFrame.
Why rolling.apply Is Extremely Slow
The key difference between rolling.mean() and your rolling.apply(partial(np.average...)) boils down to under-the-hood optimization:
rolling.mean()is built with Cython (low-level, compiled code) directly in pandas. It processes entire windows in bulk with almost no Python overhead.rolling.apply()runs a Python function (likenp.average) on every single window individually. For a 100M-row DataFrame and window sizes up to 6, that's tens of millions of Python function calls—each with its own overhead (type checking, function dispatch, etc.). This overhead adds up exponentially, leading to the 15+ minute runtimes you're seeing.
Efficient Solutions for Triangular Weighted Averages
Let's cover two high-performance approaches tailored to your triangular weights (range of window size):
1. Use Numba to Compile the Weighted Average Function
Numba compiles Python functions into machine code, eliminating most of the overhead from rolling.apply. Here's how to implement it:
import pandas as pd import numpy as np from numba import jit # Compile the weighted average function with Numba @jit(nopython=True) def triangular_weighted_avg(arr): weights = np.arange(len(arr)) # Triangular weights: 0,1,...,n-1 return (arr * weights).sum() / weights.sum() # Precompute the diff once (avoids redundant calculations) diff_mid = df['mid'].diff(1) window_sizes = [1,2,3,4,5,6] for window in window_sizes: # Use raw=True to pass numpy arrays directly (faster) df[f'triangle_mv_{window}'] = diff_mid.rolling(window).apply(triangular_weighted_avg, raw=True)
This should cut your runtime drastically—Numba-compiled functions run almost as fast as native pandas operations.
2. Use Convolution (Even Faster for Large Datasets)
Since weighted averages are essentially a convolution operation, we can leverage np.convolve (a highly optimized C implementation):
diff_mid = df['mid'].diff(1).dropna() # Remove the initial NaN from diff window_sizes = [1,2,3,4,5,6] for window in window_sizes: weights = np.arange(window) weight_sum = weights.sum() # Convolution requires reversing weights (since it's cross-correlation by default) weighted_sums = np.convolve(diff_mid, weights[::-1], mode='valid') weighted_avgs = weighted_sums / weight_sum # Pad with NaNs to align with original DataFrame length padded_avgs = pd.Series( [np.nan]*(window-1) + list(weighted_avgs), index=df.index ) df[f'triangle_mv_{window}'] = padded_avgs
Convolution is ideal for massive datasets because it processes the entire array in one pass, with no per-window Python calls.
Fixing the Custom weighted_mean Method
Your custom weighted_mean had issues with how weights were aligned to windows. Here's the corrected implementation that works as expected:
import pandas as pd import numpy as np from pandas.core.window.rolling import _Rolling_and_Expanding def weighted_mean(self, weights, **kwargs): # Convert weights to numpy array and validate length if isinstance(weights, pd.Series): weights = weights.values assert len(weights) == self.window, "Weights length must match window size" weight_sum = weights.sum() if weight_sum == 0: raise ValueError("Weights cannot sum to zero") # Define the weighted mean calculation def _calc_weighted_mean(x): return (x * weights).sum() / weight_sum # Use raw=True for faster numpy array processing return self.apply(_calc_weighted_mean, raw=True, **kwargs) # Attach the method to Rolling/Expanding objects _Rolling_and_Expanding.weighted_mean = weighted_mean # Test it out df_test = pd.DataFrame(np.reshape(range(25), (5,5))) print(df_test[1].rolling(2).weighted_mean([1,2]))
This will output the expected result:
0 NaN 1 4.333333 2 9.333333 3 14.333333 4 19.333333 Name: 1, dtype: float64
Final Recommendations
- For your 100M-row DataFrame, the convolution approach will be the fastest.
- If you need flexibility for different weight types, the Numba-compiled apply is a great balance of speed and versatility.
- Avoid raw
rolling.applywith Python functions for large datasets—always use optimized alternatives like these.
内容的提问来源于stack exchange,提问作者nick

