如何在Pandas中按列分组并标记组内指定列的最大值
问题描述
我有如下DataFrame:
| a | b | c |
|---|---|---|
| one | 6 | 11 |
| one | 7 | 12 |
| two | 8 | 23 |
| two | 9 | 14 |
| three | 10 | 15 |
| three | 20 | 25 |
我想要对列a执行groupby操作,找出每个组内列c的最大值,并为最大值所在行添加标记,最终得到如下结果:
| a | b | c | flag |
|---|---|---|---|
| one | 6 | 11 | no |
| one | 7 | 12 | yes |
| two | 8 | 23 | yes |
| two | 9 | 14 | no |
| three | 10 | 15 | no |
| three | 20 | 25 | yes |
输入DataFrame的代码如下:
import pandas as pd df = pd.DataFrame({ 'a':["one","one","two","two","three","three"], 'b':[6,7,8,9,10,20], 'c':[11,12,23,14,15,25] # , 'flag': ['no', 'yes', 'yes', 'no', 'no', 'yes'] })
解决方案
通过groupby结合transform方法获取每个组内c列的最大值,再将每行c值与组内最大值对比生成标记列:
# 计算每个分组内c列的最大值,保持与原DataFrame相同长度 max_c_per_group = df.groupby('a')['c'].transform('max') # 生成flag列:等于最大值标记为yes,否则为no df['flag'] = df['c'].eq(max_c_per_group).map({True: 'yes', False: 'no'}) print(df)
输出结果
执行代码后输出如下:
| a | b | c | flag |
|---|---|---|---|
| one | 6 | 11 | no |
| one | 7 | 12 | yes |
| two | 8 | 23 | yes |
| two | 9 | 14 | no |
| three | 10 | 15 | no |
| three | 20 | 25 | yes |
内容的提问来源于stack exchange,提问作者fast_crawler
相关产品推荐
相关产品推荐

