在Pandas中基于跨多行的两列条件生成新变量col3
Pandas 按分组条件添加二元列 col3 的实现
需求说明
给DataFrame新增一列col3,取值为yes/no,规则如下:
- 按
col1分组,若该组内所有col2的值都是yes,则组内所有行的col3为yes - 若组内
col2存在至少一个no,则组内所有行的col3为no
示例数据
import pandas as pd # 构造示例DataFrame df = pd.DataFrame({ "col1": [1,1,1,2,3,3,4,4], "col2": ["yes","no","yes","no","yes","yes","yes","no"] })
原始数据输出:
col1 col2 0 1 yes 1 1 no 2 1 yes 3 2 no 4 3 yes 5 3 yes 6 4 yes 7 4 no
实现方法
方法一:使用groupby.transform(推荐)
transform可以将分组计算的结果直接广播到原DataFrame的每一行,保证行数匹配:
# 按col1分组,判断每组所有col2是否为yes,映射为yes/no df['col3'] = df.groupby('col1')['col2'].transform(lambda x: 'yes' if (x == 'yes').all() else 'no')
方法二:先计算分组标记再映射
先得到每个col1对应的结果,再通过map匹配到原数据:
# 计算每个col1对应的标记 col1_mapping = df.groupby('col1')['col2'].apply(lambda x: 'yes' if (x == 'yes').all() else 'no') # 映射到原DataFrame df['col3'] = df['col1'].map(col1_mapping)
最终结果
运行上述代码后,得到的DataFrame如下:
col1 col2 col3 0 1 yes no 1 1 no no 2 1 yes no 3 2 no no 4 3 yes yes 5 3 yes yes 6 4 yes no 7 4 no no
内容的提问来源于stack exchange,提问作者Marco Liedecke
相关产品推荐
相关产品推荐

