You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas分组nlargest结果:将MultiIndex二级索引替换为word2值

问题:替换Pandas多级索引的二级索引值

原始DataFrame如下:

word1    word2  distance
   mango      ola        25
   mango  johnkoo        33
   apple      ola        25
   apple  johnkoo         0

我通过以下代码按word1分组,提取每组distance的前两大值:

res = df.groupby(['word1'])['distance'].nlargest(2)
print(res)

输出结果为:

word1   
apple  2    25
       3     0
mango  1    33
       0    25

这是一个带有多级索引的Pandas Series,其中二级索引是原DataFrame的行位置索引。我希望将该二级索引替换为对应的word2字段值,得到如下结果:

word1   
apple  ola      25
       johnkoo   0
mango  johnkoo  33
       ola      25

执行print(res.index)得到索引信息:

MultiIndex([('apple', 2),
            ('apple', 3),
            ('mango', 1),
            ('mango', 0)],
           names=['word1', None])

我尝试使用set_levels方法,但未能解决问题。


解决方案

方法一:分组时保留word2列(更直观)

不要单独提取distance列,而是分组后对每组按distance降序排序,取前2行,再将word1和word2设为多级索引:

res = df.groupby('word1').apply(
    lambda x: x.sort_values('distance', ascending=False).head(2)
).set_index(['word1', 'word2'])['distance']
print(res)

方法二:直接替换现有Series的二级索引

利用原DataFrame的word2列,根据二级索引的行号匹配对应值,再替换索引内容:

# 获取二级索引对应的word2值
word2_values = df.loc[res.index.get_level_values(1), 'word2'].values
# 替换二级索引
res.index = res.index.set_levels(word2_values, level=1)
# 可选:给二级索引命名
res.index.names = ['word1', 'word2']
print(res)

内容的提问来源于stack exchange,提问作者moth

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 15:10:23