如何用Pandas滚动窗口筛选列值稳定在10单位内的行?
Great question! Your current approach works for small datasets, but it has two key issues: poor performance (due to Python-level loops in the custom apply function) and a potential typo (you defined abs_delta_fn but called abs_delta in rolling().apply()). Let's fix this with a much faster, vectorized solution.
Why Your Current Code Is Slow
Custom apply functions with Python loops don't leverage Pandas' optimized C-based operations. For large DataFrames, this will become a major bottleneck—each window triggers a loop through its elements, which is orders of magnitude slower than built-in rolling methods.
The Optimized Solution
Instead of checking every pair of values in the window, we can use a simpler, equivalent logic: if the difference between the maximum and minimum values in the rolling window is ≤10.0, all values are within 10 units of each other. This lets us use Pandas' built-in rolling.max() and rolling.min() (vectorized operations that run in C), which are way faster.
Here's the code:
# Calculate rolling max and min for window size 5 rolling_max = df['column'].rolling(5, min_periods=5).max() rolling_min = df['column'].rolling(5, min_periods=5).min() # Create filter column by checking if max-min ≤10 df['filter'] = (rolling_max - rolling_min) <= 10.0
If you prefer a concise one-liner (slightly less efficient than splitting into two steps, but still vastly faster than your original code):
df['filter'] = df['column'].rolling(5, min_periods=5).apply(lambda x: x.max() - x.min() <= 10.0)
Verification Against Your Sample Data
Using your sample DataFrame:
| index | column |
|---|---|
| 1 | 0 |
| 2 | 20 |
| 3 | 40 |
| 4 | 40 |
| 5 | 40 |
| 6 | 40 |
| 7 | 40 |
| 8 | 90 |
| 9 | 120 |
| 10 | 120 |
The output matches your expected result exactly:
| index | column | filter |
|---|---|---|
| 1 | 0 | NaN |
| 2 | 20 | NaN |
| 3 | 40 | NaN |
| 4 | 40 | NaN |
| 5 | 40 | False |
| 6 | 40 | False |
| 7 | 40 | True |
| 8 | 90 | False |
| 9 | 120 | False |
| 10 | 120 | False |
Bonus: Fixing the Typo
In your original code, you defined def abs_delta_fn(window): but called df['column'].rolling(...).apply(abs_delta). This would throw a NameError—make sure the function name in apply() matches the defined function name if you ever need to use the original approach.
内容的提问来源于stack exchange,提问作者guilhermecgs

