You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas滚动窗口筛选列值稳定在10单位内的行?

Optimizing Rolling Window Stability Filter in Pandas

Great question! Your current approach works for small datasets, but it has two key issues: poor performance (due to Python-level loops in the custom apply function) and a potential typo (you defined abs_delta_fn but called abs_delta in rolling().apply()). Let's fix this with a much faster, vectorized solution.

Why Your Current Code Is Slow

Custom apply functions with Python loops don't leverage Pandas' optimized C-based operations. For large DataFrames, this will become a major bottleneck—each window triggers a loop through its elements, which is orders of magnitude slower than built-in rolling methods.

The Optimized Solution

Instead of checking every pair of values in the window, we can use a simpler, equivalent logic: if the difference between the maximum and minimum values in the rolling window is ≤10.0, all values are within 10 units of each other. This lets us use Pandas' built-in rolling.max() and rolling.min() (vectorized operations that run in C), which are way faster.

Here's the code:

# Calculate rolling max and min for window size 5
rolling_max = df['column'].rolling(5, min_periods=5).max()
rolling_min = df['column'].rolling(5, min_periods=5).min()

# Create filter column by checking if max-min ≤10
df['filter'] = (rolling_max - rolling_min) <= 10.0

If you prefer a concise one-liner (slightly less efficient than splitting into two steps, but still vastly faster than your original code):

df['filter'] = df['column'].rolling(5, min_periods=5).apply(lambda x: x.max() - x.min() <= 10.0)

Verification Against Your Sample Data

Using your sample DataFrame:

indexcolumn
10
220
340
440
540
640
740
890
9120
10120

The output matches your expected result exactly:

indexcolumnfilter
10NaN
220NaN
340NaN
440NaN
540False
640False
740True
890False
9120False
10120False

Bonus: Fixing the Typo

In your original code, you defined def abs_delta_fn(window): but called df['column'].rolling(...).apply(abs_delta). This would throw a NameError—make sure the function name in apply() matches the defined function name if you ever need to use the original approach.

内容的提问来源于stack exchange,提问作者guilhermecgs

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:51:07