如何将带多级索引的Pandas DataFrame转换为分组的嵌套字典列表
实现方法
你可以通过pandas的分组+字典转换方法直接实现,步骤和代码如下:
首先导入依赖并构造示例DataFrame:
import pandas as pd df = pd.DataFrame({ 'Index1': ['i11', 'i11', 'i21'], 'Index2': ['i12', 'i12', 'i22'], 'Index3': ['i13', 'i23', 'i23'], 'column1': [2, 3, 4], 'column2': [5, 8, 5] })
核心转换代码:
result = [] # 按Index1、Index2两列的值分组 for _, group in df.groupby(['Index1', 'Index2']): # 每组取指定列、去重后转换为每行对应一个字典的列表 group_list = group[['Index3', 'column1', 'column2']].drop_duplicates().to_dict('records') result.append(group_list)
输出result就能得到你要的嵌套结构:
[[{'Index3': 'i13', 'column1': 2, 'column2': 5}, {'Index3': 'i23', 'column1': 3, 'column2': 8}], [{'Index3': 'i23', 'column1': 4, 'column2': 5}]]
代码说明
groupby(['Index1', 'Index2'])会自动把两列取值相同的行归为同一组,符合分组规则drop_duplicates()用于过滤组内完全重复的行,满足值无重复的要求to_dict('records')是pandas内置方法,直接把DataFrame的每行转为字典,所有行组成列表,不需要手动遍历拼接
内容的提问来源于stack exchange,提问作者learner
相关产品推荐
相关产品推荐

