Python DataFrame的df.loc多列判断如何用循环替代多行重复代码
你可以先定义需要校验的最大slope序号变量,再通过逻辑条件累加的方式替代手写的重复判断,两种常用实现方式如下:
方法1:循环拼接条件
# 可按需修改这个值,比如要校验到96就设为96,校验到16就设为16 max_slope_serial = 96 df = dfm.set_index(['token', 'serial']).unstack() # 初始化第一个条件:最大序号的slope值大于0 cond = df[('slope', max_slope_serial)].gt(0) # 循环拼接后续的递减判断条件 for i in range(max_slope_serial - 1, 0, -1): cond &= df[('slope', i)].gt(df[('slope', i + 1)]) dfs = (df.loc[cond] .stack(level=1) .reset_index() .query('serial <= 1') )
方法2:用reduce批量合并条件(写法更简洁)
from functools import reduce import operator max_slope_serial = 96 df = dfm.set_index(['token', 'serial']).unstack() # 生成所有条件的列表 cond_list = [df[('slope', max_slope_serial)].gt(0)] + \ [df[('slope', i)].gt(df[('slope', i+1)]) for i in range(max_slope_serial-1, 0, -1)] # 按与逻辑合并所有条件 cond = reduce(operator.and_, cond_list) dfs = (df.loc[cond] .stack(level=1) .reset_index() .query('serial <= 1') )
两种方法执行效果完全一致,你只要修改max_slope_serial的取值就能适配任意校验长度,不需要再手动编写每一行判断逻辑。使用前请确认你的df中已经生成了序号1到max_slope_serial对应的slope列,避免索引报错。
内容的提问来源于stack exchange,提问作者teneji
相关产品推荐
相关产品推荐

