如何基于多个if条件为Pandas DataFrame创建新计算列
Pandas按多条件生成新列实现方案
实现思路
需要对三个条件分别判断,满足对应条件就累加对应分值到新列,可通过pandas条件定位或者numpy向量化判断实现,可读性和运行效率都较高。
完整代码示例
import pandas as pd # 1. 构造示例DataFrame raw_data = [ ["First", 111, 121, 323], ["Second", 222, 212, 232] ] df = pd.DataFrame( raw_data, columns=["A header", "header. 1", "header. 2", "header. 3"] ) # 2. 计算condition_head列 # 初始化分值为0 df["condition_head"] = 0 # 满足header1大于0加2 df.loc[df["header. 1"] > 0, "condition_head"] += 2 # 满足header2在2到3之间(闭区间)加5 df.loc[(df["header. 2"] >= 2) & (df["header. 2"] <= 3), "condition_head"] += 5 # 满足header3小于400加11 df.loc[df["header. 3"] < 400, "condition_head"] += 11 # 输出结果 print(df)
补充说明
按给定条件测试,两行数据的header2都远大于3,因此都不满足第二个条件,最终两行的condition_head计算结果均为13。如果需要得到示例中的目标结果,可自行调整第二个或第三个条件的数值边界即可。
如果偏好更简洁的写法,也可以用numpy的where方法实现:
import numpy as np df["condition_head"] = ( np.where(df["header. 1"] > 0, 2, 0) + np.where((df["header. 2"] >=2) & (df["header. 2"] <=3), 5, 0) + np.where(df["header. 3"] < 400, 11, 0) )
内容的提问来源于stack exchange,提问作者Archi
相关产品推荐
相关产品推荐

