Python DataFrame按Indicator与Country分组计算Value对应Rank的实现方案求助
实现方案
核心逻辑是按Indicator、Country分组后,对每组的Value列统计每个值小于等于组内所有元素的次数,两种实现方式如下:
1. 完整实现代码
方式1:使用rank方法(性能更高,适合大数据量)
pandas内置的rank方法配合method='max'参数,刚好可以直接返回小于等于当前值的元素总数,无需手动遍历:
import pandas as pd # 构造示例DataFrame data = [ { "Indicator": "A", "Country": "x", "Value": 20 }, { "Indicator": "A", "Country": "x","Value": 20 }, { "Indicator": "A","Country": "x", "Value": 30 }, { "Indicator": "B","Country": "x", "Value": 10 }, { "Indicator": "B","Country": "y","Value": 30 }, { "Indicator": "B", "Country": "y", "Value": 20 } ] df = pd.DataFrame(data) # 计算Rank列 df['Rank'] = df.groupby(['Indicator', 'Country'])['Value']\ .transform(lambda x: x.rank(method='max', ascending=True))\ .astype(int)
方式2:手动计数(逻辑更直观,无理解门槛)
完全贴合需求描述的原生逻辑实现,直接遍历组内每个值统计符合条件的次数:
df['Rank'] = df.groupby(['Indicator', 'Country'])['Value'].transform( lambda ser: [sum(ser <= val) for val in ser] )
2. 输出结果
运行上述代码后最终得到的DataFrame如下:
| Indicator | Country | Value | Rank |
|---|---|---|---|
| A | x | 20 | 2 |
| A | x | 20 | 2 |
| A | x | 30 | 3 |
| B | x | 10 | 1 |
| B | y | 30 | 2 |
| B | y | 20 | 1 |
内容的提问来源于stack exchange,提问作者MLotz
相关产品推荐
相关产品推荐

