Pandas分组nlargest结果:将MultiIndex二级索引替换为word2值
问题:替换Pandas多级索引的二级索引值
原始DataFrame如下:
word1 word2 distance mango ola 25 mango johnkoo 33 apple ola 25 apple johnkoo 0
我通过以下代码按word1分组,提取每组distance的前两大值:
res = df.groupby(['word1'])['distance'].nlargest(2) print(res)
输出结果为:
word1 apple 2 25 3 0 mango 1 33 0 25
这是一个带有多级索引的Pandas Series,其中二级索引是原DataFrame的行位置索引。我希望将该二级索引替换为对应的word2字段值,得到如下结果:
word1 apple ola 25 johnkoo 0 mango johnkoo 33 ola 25
执行print(res.index)得到索引信息:
MultiIndex([('apple', 2), ('apple', 3), ('mango', 1), ('mango', 0)], names=['word1', None])
我尝试使用set_levels方法,但未能解决问题。
解决方案
方法一:分组时保留word2列(更直观)
不要单独提取distance列,而是分组后对每组按distance降序排序,取前2行,再将word1和word2设为多级索引:
res = df.groupby('word1').apply( lambda x: x.sort_values('distance', ascending=False).head(2) ).set_index(['word1', 'word2'])['distance'] print(res)
方法二:直接替换现有Series的二级索引
利用原DataFrame的word2列,根据二级索引的行号匹配对应值,再替换索引内容:
# 获取二级索引对应的word2值 word2_values = df.loc[res.index.get_level_values(1), 'word2'].values # 替换二级索引 res.index = res.index.set_levels(word2_values, level=1) # 可选:给二级索引命名 res.index.names = ['word1', 'word2'] print(res)
内容的提问来源于stack exchange,提问作者moth
相关产品推荐
相关产品推荐

