如何将多组针对DataFrame的np.where()条件合并为一行代码?
合并多组np.where()语句为单行代码
首先,你的原始DataFrame定义如下:
import numpy as np import pandas as pd a = [{'name': 'A', 'col_1': 3, 'col_2': 2, 'col_3': 0.5, 'col_4': 0.2, 'col_5': 1}, {'name': 'A', 'col_1': 1, 'col_2': 0, 'col_3': 0.5, 'col_4': 0.2, 'col_5': 1}, {'name': 'B', 'col_1': 3, 'col_2': 2, 'col_3': 2, 'col_4': 0.2, 'col_5': 2}, {'name': 'B', 'col_1': 1, 'col_2': 0, 'col_3': 0, 'col_4': 0.2, 'col_5': 2}, {'name': 'C', 'col_1': 3, 'col_2': 2, 'col_3': 0.5, 'col_4': 2, 'col_5': 3}, {'name': 'C', 'col_1': 1, 'col_2': 2, 'col_3': 0.5, 'col_4': 0, 'col_5': 3}] df = pd.DataFrame(a)
你原本通过三次np.where()赋值生成new列:
df['new'] = np.where((df['col_5'] == 1) & (df['col_2'] != 0), df['col_2'], df['col_1'] * 0.25) df['new'] = np.where((df['col_5'] == 2) & (df['col_3'] != 0), df['col_3'], df['col_1'] * 0.5) df['new'] = np.where((df['col_5'] == 3) & (df['col_4'] != 0), df['col_4'], df['col_1'] * 0.75)
合并为单行代码的方案
推荐用np.select()实现,它能一次性处理多组条件与对应值,逻辑和原代码完全一致,且可读性更强:
df['new'] = np.select( condlist=[ (df['col_5'] == 1) & (df['col_2'] != 0), df['col_5'] == 1, (df['col_5'] == 2) & (df['col_3'] != 0), df['col_5'] == 2, (df['col_5'] == 3) & (df['col_4'] != 0), df['col_5'] == 3 ], choicelist=[ df['col_2'], df['col_1'] * 0.25, df['col_3'], df['col_1'] * 0.5, df['col_4'], df['col_1'] * 0.75 ], default=df['col_1'] * 0.75 # 兜底值,实际业务中不会触发 )
也可以用嵌套np.where()实现,但结构相对繁琐:
df['new'] = np.where( df['col_5'] == 1, np.where(df['col_2'] != 0, df['col_2'], df['col_1'] * 0.25), np.where( df['col_5'] == 2, np.where(df['col_3'] != 0, df['col_3'], df['col_1'] * 0.5), np.where(df['col_4'] != 0, df['col_4'], df['col_1'] * 0.75) ) )
两种方法都能得到和原代码完全相同的结果,优先推荐np.select(),便于后续维护。
内容的提问来源于stack exchange,提问作者LiAfe
相关产品推荐
相关产品推荐

