You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何替代for循环,更高效地从嵌套字典user_dict生成指定Pandas DataFrame?

从嵌套字典生成指定结构DataFrame的更优实现

问题描述

我需要从嵌套字典user_dict生成指定结构的DataFrame,目前已经通过for循环实现了需求,想知道有没有更简洁高效的实现方式?

原实现代码

import pandas as pd

user_dict = {12: {'Category 1': {'att_1': 1, 'att_2': 'whatever'}, 
                  'Category 2': {'att_1': 23, 'att_2': 'another'}},
             15: {'Category 1': {'att_1': 10, 'att_2': 'foo'}, 
                  'Category 2': {'att_1': 30, 'att_2': 'bar'}}}

mydict = {"level1": [], "level2": [], "att_1": [], "att_2": [] }

for level1key, leve1value in user_dict.items() :
     for level2key, level2value in leve1value.items():
          mydict["level1"].append(level1key)
          mydict["level2"].append(level2key)
          for level3key, level3value in level2value.items():
             if level3key == "att_1": 
               mydict["att_1"].append(level3value)
             else:
               mydict["att_2"].append(level3value)

df = pd.DataFrame(mydict)
print(df)

目标输出

level1      level2  att_1     att_2
0      12  Category 1      1  whatever
1      12  Category 2     23   another
2      15  Category 1     10       foo
3      15  Category 2     30       bar

更优实现方式

方法1:利用pd.json_normalize(最简洁)

借助pandas内置的json_normalize处理嵌套结构,再通过melt和字典展开完成列整理:

import pandas as pd

user_dict = {12: {'Category 1': {'att_1': 1, 'att_2': 'whatever'}, 
                  'Category 2': {'att_1': 23, 'att_2': 'another'}},
             15: {'Category 1': {'att_1': 10, 'att_2': 'foo'}, 
                  'Category 2': {'att_1': 30, 'att_2': 'bar'}}}

# 转换为包含外层键的字典列表
data = [{"level1": key, **value} for key, value in user_dict.items()]
# 展开嵌套结构并重塑表格
df = pd.json_normalize(data, sep='_').melt(id_vars='level1', var_name='level2', value_name='attrs')
# 展开属性字典为列
df = pd.concat([df, df['attrs'].apply(pd.Series)], axis=1).drop('attrs', axis=1)
# 调整列顺序匹配目标结构
df = df[['level1', 'level2', 'att_1', 'att_2']].reset_index(drop=True)
print(df)

方法2:利用多层索引stack展开

通过构造多层索引的Series,再展开为DataFrame:

import pandas as pd

user_dict = {12: {'Category 1': {'att_1': 1, 'att_2': 'whatever'}, 
                  'Category 2': {'att_1': 23, 'att_2': 'another'}},
             15: {'Category 1': {'att_1': 10, 'att_2': 'foo'}, 
                  'Category 2': {'att_1': 30, 'att_2': 'bar'}}}

# 构造多层索引Series并展开
s = pd.Series(user_dict).apply(pd.Series).stack()
# 从索引提取level1和level2,同时展开属性字典
df = pd.DataFrame(s.tolist()).assign(
    level1=s.index.get_level_values(0),
    level2=s.index.get_level_values(1)
)
# 调整列顺序
df = df[['level1', 'level2', 'att_1', 'att_2']].reset_index(drop=True)
print(df)

方法3:简化版列表推导

相比原循环,直接构造每行的字典,减少多次列表append的性能损耗:

import pandas as pd

user_dict = {12: {'Category 1': {'att_1': 1, 'att_2': 'whatever'}, 
                  'Category 2': {'att_1': 23, 'att_2': 'another'}},
             15: {'Category 1': {'att_1': 10, 'att_2': 'foo'}, 
                  'Category 2': {'att_1': 30, 'att_2': 'bar'}}}

# 一次性生成所有行数据
rows = []
for level1, categories in user_dict.items():
    for cat_name, attrs in categories.items():
        rows.append({
            'level1': level1,
            'level2': cat_name,
            'att_1': attrs['att_1'],
            'att_2': attrs['att_2']
        })

df = pd.DataFrame(rows)
print(df)

优势说明

  • 方法1和2充分利用pandas的内置函数,代码更简洁,可读性更强,适合处理复杂嵌套结构;
  • 方法3相比原循环,减少了多次append操作(列表多次append会产生额外内存开销),性能更优,同时逻辑更清晰。

内容的提问来源于stack exchange,提问作者Pete

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 04:14:56