Python中基于双条件为DataFrame新增列的实现求助
解决Pandas DataFrame生成指定输出列的问题
示例数据集
import pandas as pd dict_1 = {'Id' : [1, 1, 2, 2, 3, 4], 'boolean_val' : [True, False, True, False, True, False], "sal" : [1000, 2000, 1500, 2500, 3500, 4500]} test = pd.DataFrame(dict_1)
需求规则
- 新增
output_True列:- 若同一
Id存在两行且当前行boolean_val为True,填入该行sal值;否则填"NA"
- 若同一
- 新增
output_False列:- 若同一
Id存在两行且当前行boolean_val为False,填入该行sal值;否则填"NA"
- 若同一
- 特殊场景:若
Id唯一(仅一行),则boolean_val为True时将sal填入output_True,为False时填入output_False
期望输出
expected_dict = {'Id' : [1, 1, 2, 2, 3, 4], 'boolean_val' : [True, False, True, False, True, False], "sal" : [1000, 2000, 1500, 2500, 3500, 4500], "output_True" : [1000, "NA", 1500, "NA", 3500, "NA"], "output_False" : [2000, "NA", 2500, "NA", "NA", 4500]} output_df = pd.DataFrame(expected_dict)
解决方案
核心思路是先标记每个Id的行数,再结合boolean_val的条件填充目标列:
import pandas as pd import numpy as np # 加载数据集 dict_1 = {'Id' : [1, 1, 2, 2, 3, 4], 'boolean_val' : [True, False, True, False, True, False], "sal" : [1000, 2000, 1500, 2500, 3500, 4500]} test = pd.DataFrame(dict_1) # 生成辅助列:每个Id对应的行数 test['id_count'] = test.groupby('Id')['Id'].transform('count') # 填充output_True列 test['output_True'] = np.where( ((test['id_count'] == 2) & (test['boolean_val'] == True)) | ((test['id_count'] == 1) & (test['boolean_val'] == True)), test['sal'], 'NA' ) # 填充output_False列 test['output_False'] = np.where( ((test['id_count'] == 2) & (test['boolean_val'] == False)) | ((test['id_count'] == 1) & (test['boolean_val'] == False)), test['sal'], 'NA' ) # 移除辅助列(可选) test = test.drop('id_count', axis=1) # 查看结果 print(test)
代码说明
- 标记Id行数:通过
groupby+transform生成id_count列,快速区分当前行的Id是成对出现还是唯一。 - 条件填充:用
np.where组合两个判断条件,精准匹配需求中的场景,满足条件则填入sal值,否则填"NA"。 - 清理数据:如果不需要辅助列,可通过
drop删除。
运行代码后,输出结果与期望完全一致。
内容的提问来源于stack exchange,提问作者Data-7scientist
相关产品推荐
相关产品推荐

