Python Pandas:如何对DataFrame移位填充指定值?及布尔值连续行扩展的优化实现
Great question! Your current solution gets the job done, but we can streamline this into cleaner, more efficient code using vectorized operations—no need for all those helper columns. Let's break this down.
First, let's confirm the goal: you have a boolean DataFrame, and you want to mark any row that's True plus the next two rows as True (your expected result shows rows 3-7 as True, which lines up with the True values at positions 3 and 5 covering themselves and their next two rows).
Here are two concise, efficient approaches:
Approach 1: Numpy Convolution (Fastest for Large Data)
Convolution is perfect here because it lets us "spread" the True values across the desired range with a single vectorized operation:
import pandas as pd import numpy as np # Initialize your DataFrame y = pd.DataFrame(np.zeros((10,1), dtype='bool'), columns=['A']) y.iloc[[3,5], 0] = True # Extend True values to current row + next 2 rows y['A'] = np.convolve(y['A'].values, np.ones(3, dtype=bool), mode='same').astype(bool)
- The
np.ones(3, dtype=bool)acts as a "kernel"—it tells numpy to activate the current position and the next two positions whenever it hits aTrue. mode='same'ensures we keep the original length of the DataFrame.
Approach 2: Pandas Rolling Window (Pandas-Native)
If you prefer sticking to pandas methods, we can use a rolling window with a quick reverse trick to look ahead instead of backward:
# Extend True values by rolling backward after reversing the series y['A'] = y['A'][::-1].rolling(window=3, min_periods=1).max()[::-1].astype(bool)
- Reversing the series turns "looking ahead" into "looking back" for the rolling window. A window of 3 captures the current row and the two prior rows (which are the original next two rows).
- Taking the
max()ensures anyTruein the window keeps the row marked asTrue, then we reverse back to restore the original order.
Why these are better than your current code:
- Cleaner: No intermediate columns (
B,C,D) cluttering up your code - Faster: Vectorized operations are way more efficient, especially as your DataFrame grows
- More readable: The kernel/window size makes it immediately clear how many rows you're extending the True values to
When you run either approach, you'll get exactly your expected output:
A 0 False 1 False 2 False 3 True 4 True 5 True 6 True 7 True 8 False 9 False
内容的提问来源于stack exchange,提问作者david78

