You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas DataFrame中获取各唯一行序列的对应索引列表?

问题:将Pandas DataFrame的唯一行映射到对应的行索引列表

现有如下Pandas DataFrame(列名默认是[0,1,2,3,4]):

import pandas as pd
df = pd.DataFrame(  [[0,0,1,0,0],
                    [0,0,0,0,1],
                    [0,0,1,0,0],
                    [0,0,0,0,1],
                    [0,0,1,0,0],
                    [0,0,0,0,1],
                    [0,0,1,0,0],
                    [0,0,0,1,0],
                    [0,1,1,0,0],
                    [1,0,0,1,0],
                    [0,1,1,0,0],
                    [1,0,0,1,0],
                    [0,1,1,0,1]])

运行df.value_counts()可得到唯一行序列及其出现次数,但需要生成一个字典,其中键为各唯一行序列,值为该序列对应的所有行索引列表,示例输出如下(已修正原示例中的索引错误):

{(0,0,1,0,0): [0,2,4,6],
 (0,0,0,0,1): [1,3,5],
 (0,1,1,0,0): [8,10],
 (1,0,0,1,0): [9,11],
 (0,0,0,1,0): [7],
 (0,1,1,0,1): [12]}
解决方案

可以通过Pandas的groupby结合元组转换实现,具体步骤如下:

  1. 将每行数据转换为元组(列表不可哈希,无法作为字典键)
  2. 按行元组分组,收集每组的索引并转换为列表
  3. 将分组结果转换为目标字典

代码实现:

# 按行元组分组,获取每组的索引数组
grouped_indices = df.groupby(df.apply(tuple, axis=1)).indices

# 将索引数组转为列表,生成目标字典
result = {row_tuple: indices.tolist() for row_tuple, indices in grouped_indices.items()}

print(result)

输出结果:

{(0, 0, 0, 0, 1): [1, 3, 5],
 (0, 0, 0, 1, 0): [7],
 (0, 0, 1, 0, 0): [0, 2, 4, 6],
 (0, 1, 1, 0, 0): [8, 10],
 (0, 1, 1, 0, 1): [12],
 (1, 0, 0, 1, 0): [9, 11]}

补充说明:

  • df.apply(tuple, axis=1):逐行将DataFrame的行数据转换为元组,作为分组的依据
  • groupby(...).indices:返回字典,键为分组的行元组,值为对应行索引的numpy数组
  • 字典推导式将numpy数组转为普通列表,得到最终需要的格式

内容的提问来源于stack exchange,提问作者Saran Pannasuriyaporn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 22:01:17