Pandas中根据列值选择不同聚合函数的单行代码实现
单行代码实现DataFrame按标签动态生成聚合列
原始数据
先定义测试用的DataFrame:
import pandas as pd import numpy as np test = pd.DataFrame({'y':[1,2,3,4,5,6], 'label': ['bottom', 'top','bottom', 'top','bottom', 'top']})
对应输出:
y label 0 1 bottom 1 2 top 2 3 bottom 3 4 top 4 5 bottom 5 6 top
需求
新增agg_y列:
- 当
label为bottom时,取该分组下y的最大值 - 当
label为top时,取该分组下y的最小值
单行解决方案
提供两种简洁的单行实现方式:
方式1:np.where结合分组变换
直接通过np.where分支调用对应的分组聚合变换:
test['agg_y'] = np.where(test['label'] == 'bottom', test.groupby('label')['y'].transform('max'), test.groupby('label')['y'].transform('min'))
方式2:自定义分组聚合逻辑
利用分组名(即label值)映射对应的聚合操作,代码更紧凑:
test['agg_y'] = test.groupby('label')['y'].transform(lambda g: {'bottom': g.max(), 'top': g.min()}[g.name])
执行结果
两种方式均能得到预期输出:
y label agg_y 0 1 bottom 5 1 2 top 2 2 3 bottom 5 3 4 top 2 4 5 bottom 5 5 6 top 2
内容的提问来源于stack exchange,提问作者quant
相关产品推荐
相关产品推荐

