Pandas DataFrame 指定列遇1时将后续5行置为0的实现方法
最优实现方案(numpy向量化,避免全量逐行迭代)
import pandas as pd import numpy as np # 构造原始DataFrame df = pd.DataFrame({'Col1':[0,1,0,1,0,1,1,0,1,0,0,0,1,0,0,1,0,1,0,0,0]}) arr = df['Col1'].values n = len(arr) # 找到所有值为1的索引位置 ones_idx = np.where(arr == 1)[0] # 初始化掩码:标记需要置为0的位置 mask = np.zeros(n, dtype=bool) for idx in ones_idx: # 仅对当前1的位置之后连续5行标记为需要置0 mask[idx+1 : min(idx+6, n)] = True # 生成处理后的结果列 df['Col1'] = np.where(mask, 0, arr)
运行后输出和你给出的预期结果完全一致:
>>> df['Col1'].tolist() [0, 1, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0]
方案说明
- 仅循环所有值为1的索引,而非逐行遍历全表,数据量越大相比逐行迭代的性能优势越明显
- 自动处理边界场景:如果1出现在最后5行以内,后续不足5行的部分也会自动全部置0
如果是百万行以上的超大数据量,可以用滑动窗口实现完全向量化,进一步提升性能:
from numpy.lib.stride_tricks import sliding_window_view arr = df['Col1'].values n = len(arr) if n >= 6: # 构造滑动窗口判断每个位置前5行是否存在1 windows = sliding_window_view(np.pad(arr, (5, 0), constant_values=0), 5)[:-1] mask = (windows == 1).any(axis=1) else: mask = np.zeros(n, dtype=bool) df['Col1'] = np.where(mask, 0, arr)
内容的提问来源于stack exchange,提问作者Felton Wang
相关产品推荐
相关产品推荐

