Python如何在分组数据表中新增列展示组内排名前两位的变量值
实现方案
前置优化(可选)
你当前的排名代码默认采用method='average'参数,若存在相同数值会生成小数排名,建议先调整为整数不重复排名,方便后续取Top值:
df['Rank'] = df.groupby('Group')['Value'].rank(ascending=True, method='first').astype(int)
新增分组Top1、Top2数值列
方法1:不依赖已生成的Rank列,直接计算
import pandas as pd # 分组内排名第一的数值(升序排名第一即为组内最小值) df['Group_Top1_Value'] = df.groupby('Group')['Value'].transform('min') # 分组内排名第二的数值,分组不足2条时返回空值 df['Group_Top2_Value'] = df.groupby('Group')['Value'].transform( lambda x: x.sort_values().iloc[1] if len(x) >= 2 else pd.NA )
方法2:基于已有的Rank列计算,和你当前的排名规则完全对齐
# 取每组排名为1的数值,广播到组内所有行 df['Group_Top1_Value'] = df[df['Rank'] == 1].groupby('Group')['Value'].first().reindex(df['Group']).values # 取每组排名为2的数值,广播到组内所有行 df['Group_Top2_Value'] = df[df['Rank'] == 2].groupby('Group')['Value'].first().reindex(df['Group']).values
注意事项
如果存在并列排名需求,可调整rank函数的method参数:
method='min':相同数值取最小排名,会出现多个同排名条目method='dense':排名连续,相同数值排名一致,后续排名不跳号
内容的提问来源于stack exchange,提问作者Richard Dixon
相关产品推荐
相关产品推荐

