Python基于多条件规则为pandas DataFrame新增result列实现求助
pandas DataFrame按多条件规则新增result列实现方案
方案1:使用numpy.select(推荐,大体积数据效率更高)
该方法是向量化操作,相比逐行遍历的apply方法性能优势明显,适合绝大多数场景。
完整实现代码:
import pandas as pd import numpy as np # 示例原始数据(实际使用时替换成你自己的DataFrame即可) data = [['1', 'stp', 'salaried', '> 10 lakh'], ['2', 'stp', 'business', '> 10 lakh'], ['3', 'n_stp', 'salaried', '<= 10 lakh'], ['4', 'n_stp', 'other', '任意值'], ['5', 'stp', 'other', '任意值']] df = pd.DataFrame(data, columns = ['s.no', 'flg', 'emp', 'inc']) # 按规则定义条件列表和对应返回值列表,顺序要一一对应 conditions = [ # 规则1 df['emp'].isin(['salaried', 'business']) & (df['inc'] == '> 10 lakh') & (df['flg'] == 'stp'), # 规则2 df['emp'].isin(['salaried', 'business']) & (df['inc'] == '<= 10 lakh') & (df['flg'] == 'n_stp'), # 规则5 (df['emp'] == 'other') & (df['flg'] == 'n_stp'), # 规则3 df['emp'].isin(['salaried', 'business']) & (df['inc'] == '> 10 lakh') & (df['flg'] == 'n_stp'), # 规则4 df['emp'].isin(['salaried', 'business']) & (df['inc'] == '<= 10 lakh') & (df['flg'] == 'stp'), # 规则6 (df['emp'] == 'other') & (df['flg'] == 'stp') ] choices = [ 'no_issue', 'no_issue', 'no issue', 'issue', 'issue', 'issue' ] # 新增result列,default参数为不满足所有规则时的默认填充值,可按需修改 df['result'] = np.select(conditions, choices, default=np.nan) print(df)
运行输出结果:
s.no flg emp inc result 0 1 stp salaried > 10 lakh no_issue 1 2 stp business > 10 lakh no_issue 2 3 n_stp salaried <= 10 lakh no_issue 3 4 n_stp other 任意值 no issue 4 5 stp other 任意值 issue
方案2:使用自定义函数+apply(适合逻辑后续要频繁调整的场景)
如果后续规则还会经常迭代,自定义函数的可读性和可维护性更高,缺点是数据量超过10万行时性能较低。
实现代码:
import pandas as pd # 示例数据 data = [['1', 'stp', 'salaried', '> 10 lakh'], ['2', 'stp', 'business', '> 10 lakh'], ['3', 'n_stp', 'salaried', '<= 10 lakh']] df = pd.DataFrame(data, columns = ['s.no', 'flg', 'emp', 'inc']) # 自定义规则匹配函数 def match_rule(row): emp = row['emp'] flg = row['flg'] inc = row['inc'] if emp in ['salaried', 'business']: if (inc == '> 10 lakh' and flg == 'stp') or (inc == '<= 10 lakh' and flg == 'n_stp'): return 'no_issue' else: return 'issue' elif emp == 'other': return 'no issue' if flg == 'n_stp' else 'issue' # 可补充其他emp取值的规则,默认返回None return None # 逐行应用函数生成新列 df['result'] = df.apply(match_rule, axis=1)
内容的提问来源于stack exchange,提问作者KReEd
相关产品推荐
相关产品推荐

