如何更简洁地将DataFrame转换为多级嵌套字典?
Pandas DataFrame转多级嵌套字典(替代双重for循环)
需求背景
已有如下DataFrame,通过双重for循环实现了转多级嵌套字典,但希望用更简洁的方式实现,避免嵌套循环:
from pprint import pprint import pandas as pd a = pd.DataFrame([ {'col1': 'A', 'col2': 'Person 1', 'height': 1, 'weight': 10}, {'col1': 'A', 'col2': 'Person 1', 'height': 2, 'weight': 20}, {'col1': 'A', 'col2': 'Person 1', 'height': 3, 'weight': 30}, {'col1': 'A', 'col2': 'Person 2', 'height': 4, 'weight': 40}, {'col1': 'A', 'col2': 'Person 2', 'height': 5, 'weight': 50}, {'col1': 'A', 'col2': 'Person 2', 'height': 6, 'weight': 60}, {'col1': 'B', 'col2': 'Person 1', 'height': 11, 'weight': 101}, {'col1': 'B', 'col2': 'Person 1', 'height': 21, 'weight': 201}, {'col1': 'B', 'col2': 'Person 1', 'height': 31, 'weight': 301}, {'col1': 'B', 'col2': 'Person 2', 'height': 41, 'weight': 401}, {'col1': 'B', 'col2': 'Person 2', 'height': 51, 'weight': 501}, {'col1': 'B', 'col2': 'Person 2', 'height': 61, 'weight': 601}, ]) # 原始双重循环实现 result = {} for col1, j in a.groupby('col1'): result[col1] = {} for col2, n in j.groupby('col2'): result[col1][col2] = n.to_dict(orient='records') pprint(result)
预期结果
{'A': {'Person 1': [{'col1': 'A', 'col2': 'Person 1', 'height': 1, 'weight': 10}, {'col1': 'A', 'col2': 'Person 1', 'height': 2, 'weight': 20}, {'col1': 'A', 'col2': 'Person 1', 'height': 3, 'weight': 30}], 'Person 2': [{'col1': 'A', 'col2': 'Person 2', 'height': 4, 'weight': 40}, {'col1': 'A', 'col2': 'Person 2', 'height': 5, 'weight': 50}, {'col1': 'A', 'col2': 'Person 2', 'height': 6, 'weight': 60}]}, 'B': {'Person 1': [{'col1': 'B', 'col2': 'Person 1', 'height': 11, 'weight': 101}, {'col1': 'B', 'col2': 'Person 1', 'height': 21, 'weight': 201}, {'col1': 'B', 'col2': 'Person 1', 'height': 31, 'weight': 301}], 'Person 2': [{'col1': 'B', 'col2': 'Person 2', 'height': 41, 'weight': 401}, {'col1': 'B', 'col2': 'Person 2', 'height': 51, 'weight': 501}, {'col1': 'B', 'col2': 'Person 2', 'height': 61, 'weight': 601}]}}
问题:尝试的代码未正确嵌套col2
之前尝试的代码无法生成预期的嵌套结构:
a.groupby('col1').apply(lambda x: x.set_index('col2').to_dict(orient='records')).to_dict()
简洁解决方案
方法1:双层groupby + unstack + to_dict
通过同时按col1和col2分组,将每组转为records列表,再通过unstack把col2转为列维度,最后转为字典:
result = a.groupby(['col1', 'col2']).apply(lambda x: x.to_dict('records')).unstack().to_dict('index')
方法2:嵌套groupby + apply
在col1分组后的每个子组内,再按col2分组并转为字典:
result = a.groupby('col1').apply( lambda g: g.groupby('col2').apply(lambda x: x.to_dict('records')).to_dict() ).to_dict()
方法3:使用字典推导式(替代显式循环)
如果不想用嵌套apply,也可以用字典推导式简化循环逻辑,比原始双重循环更简洁:
result = { col1: {col2: df.to_dict('records') for col2, df in g.groupby('col2')} for col1, g in a.groupby('col1') }
以上三种方法均可生成与原始双重循环完全一致的结果,且代码更简洁。
内容的提问来源于stack exchange,提问作者Simon1
相关产品推荐
相关产品推荐

