You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何更简洁地将DataFrame转换为多级嵌套字典?

Pandas DataFrame转多级嵌套字典(替代双重for循环)

需求背景

已有如下DataFrame,通过双重for循环实现了转多级嵌套字典,但希望用更简洁的方式实现,避免嵌套循环:

from pprint import pprint
import pandas as pd

a = pd.DataFrame([
    {'col1': 'A', 'col2': 'Person 1', 'height': 1, 'weight': 10},
    {'col1': 'A', 'col2': 'Person 1', 'height': 2, 'weight': 20},
    {'col1': 'A', 'col2': 'Person 1', 'height': 3, 'weight': 30},
    {'col1': 'A', 'col2': 'Person 2', 'height': 4, 'weight': 40},
    {'col1': 'A', 'col2': 'Person 2', 'height': 5, 'weight': 50},
    {'col1': 'A', 'col2': 'Person 2', 'height': 6, 'weight': 60},
    {'col1': 'B', 'col2': 'Person 1', 'height': 11, 'weight': 101},
    {'col1': 'B', 'col2': 'Person 1', 'height': 21, 'weight': 201},
    {'col1': 'B', 'col2': 'Person 1', 'height': 31, 'weight': 301},
    {'col1': 'B', 'col2': 'Person 2', 'height': 41, 'weight': 401},
    {'col1': 'B', 'col2': 'Person 2', 'height': 51, 'weight': 501},
    {'col1': 'B', 'col2': 'Person 2', 'height': 61, 'weight': 601},
])

# 原始双重循环实现
result = {}
for col1, j in a.groupby('col1'):
    result[col1] = {}
    for col2, n in j.groupby('col2'):
        result[col1][col2] = n.to_dict(orient='records')

pprint(result)

预期结果

{'A': {'Person 1': [{'col1': 'A',
                     'col2': 'Person 1',
                     'height': 1,
                     'weight': 10},
                    {'col1': 'A',
                     'col2': 'Person 1',
                     'height': 2,
                     'weight': 20},
                    {'col1': 'A',
                     'col2': 'Person 1',
                     'height': 3,
                     'weight': 30}],
       'Person 2': [{'col1': 'A',
                     'col2': 'Person 2',
                     'height': 4,
                     'weight': 40},
                    {'col1': 'A',
                     'col2': 'Person 2',
                     'height': 5,
                     'weight': 50},
                    {'col1': 'A',
                     'col2': 'Person 2',
                     'height': 6,
                     'weight': 60}]},
 'B': {'Person 1': [{'col1': 'B',
                     'col2': 'Person 1',
                     'height': 11,
                     'weight': 101},
                    {'col1': 'B',
                     'col2': 'Person 1',
                     'height': 21,
                     'weight': 201},
                    {'col1': 'B',
                     'col2': 'Person 1',
                     'height': 31,
                     'weight': 301}],
       'Person 2': [{'col1': 'B',
                     'col2': 'Person 2',
                     'height': 41,
                     'weight': 401},
                    {'col1': 'B',
                     'col2': 'Person 2',
                     'height': 51,
                     'weight': 501},
                    {'col1': 'B',
                     'col2': 'Person 2',
                     'height': 61,
                     'weight': 601}]}}

问题:尝试的代码未正确嵌套col2

之前尝试的代码无法生成预期的嵌套结构:

a.groupby('col1').apply(lambda x: x.set_index('col2').to_dict(orient='records')).to_dict()

简洁解决方案

方法1:双层groupby + unstack + to_dict

通过同时按col1和col2分组,将每组转为records列表,再通过unstack把col2转为列维度,最后转为字典:

result = a.groupby(['col1', 'col2']).apply(lambda x: x.to_dict('records')).unstack().to_dict('index')

方法2:嵌套groupby + apply

在col1分组后的每个子组内,再按col2分组并转为字典:

result = a.groupby('col1').apply(
    lambda g: g.groupby('col2').apply(lambda x: x.to_dict('records')).to_dict()
).to_dict()

方法3:使用字典推导式(替代显式循环)

如果不想用嵌套apply,也可以用字典推导式简化循环逻辑,比原始双重循环更简洁:

result = {
    col1: {col2: df.to_dict('records') for col2, df in g.groupby('col2')}
    for col1, g in a.groupby('col1')
}

以上三种方法均可生成与原始双重循环完全一致的结果,且代码更简洁。


内容的提问来源于stack exchange,提问作者Simon1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 10:22:45