基于col1的1触发条件查找col2首个1并计算两位置索引差
实现方案
核心逻辑是预先提取所有col2值为1的升序索引列表,对每个col1为1的索引用二分查找快速定位后续第一个符合要求的col2为1的索引,计算差值即可,该方法时间复杂度仅为O(n log n),适配任意长度和分布的输入数据。
完整代码实现
import pandas as pd import numpy as np from bisect import bisect_right # ---------- 1. 构造测试数据(实际使用时替换为你自己的df即可)---------- np.random.seed(42) df = pd.DataFrame({ 'col1': np.random.choice([0,1], size=20, p=[0.7, 0.3]), 'col2': np.random.choice([0,1], size=20, p=[0.7, 0.3]) }) # ---------- 2. 预处理:提取所有col2为1的索引列表 ---------- col2_one_indexes = df[df['col2'] == 1].index.tolist() # ---------- 3. 差值计算函数 ---------- def get_index_diff(current_idx): # 二分查找第一个大于当前索引的col2为1的位置 target_pos = bisect_right(col2_one_indexes, current_idx) if target_pos < len(col2_one_indexes): return col2_one_indexes[target_pos] - current_idx # 后续没有col2为1的情况,返回空值,可根据需求修改默认值 return np.nan # ---------- 4. 生成结果列 ---------- # col1不为1的位置默认返回空值,也可根据需求修改为0或其他值 df['index_diff'] = df.apply( lambda row: get_index_diff(row.name) if row['col1'] == 1 else np.nan, axis=1 ) # 打印结果验证 print(df)
补充说明
- 如果需要计算行位置差而非索引值差,只需要提前将df重置为默认连续整数索引即可:
df = df.reset_index(drop=True) - 若希望
col1不为1、或后续无col2为1的位置返回0而非空值,直接修改对应位置的返回值即可
内容的提问来源于stack exchange,提问作者Hannan Naseem
相关产品推荐
相关产品推荐

