You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将展开的DataFrame转换为保留列名层级的深度嵌套字典

如何将展开的DataFrame转换为保留列名层级的深度嵌套字典

我完全懂你的困扰——把扁平化的DataFrame转成嵌套字典时丢失了列名层级,确实会让字典结构变得混乱,根本没法轻松导航。别担心,我们可以修改递归逻辑,让每一层都明确保留列名作为父节点的键,而不是直接用列值当顶层键,这样整个结构就清晰多了。

解决方案:修改递归函数保留列名层级

下面是调整后的函数,它会在每一层嵌套中都保留对应的列名作为键,让你能清楚知道当前层级对应的是DataFrame的哪个字段:

import pandas as pd

def frame_to_nested_dict_with_columns(df, levels):
    # 递归终止条件:没有剩余层级时,构建最内层的指标-值映射
    if not levels:
        return df.groupby('enrollment_measure_name')['value'].first().to_dict()
    
    current_col = levels[0]
    nested_dict = {}
    
    # 按当前层级的列分组,遍历每个分组
    for col_value, group in df.groupby(current_col):
        # 先创建以列名为键的父节点,避免覆盖
        if current_col not in nested_dict:
            nested_dict[current_col] = {}
        # 递归处理剩余层级,将结果挂载到当前列值对应的位置
        nested_dict[current_col][col_value] = frame_to_nested_dict_with_columns(group, levels[1:])
    
    return nested_dict

调用示例

用你提供的exploded_df来测试,指定你想要的嵌套层级顺序(比如school_code → district_code → year → source):

# 你的原始DataFrame
exploded_df = pd.DataFrame({
    'school_code': [1, 1, 1, 1, 2, 2,  2, 2],
    'school_name': ['A', 'A', 'A', 'A', 'B', 'B', 'B', 'B'],
    'district_code': [10, 10, 10, 10, 20, 20,  20, 20],
    'year': [2022, 2022, 2023, 2023, 2022, 2022, 2023, 2023],
    'source': ['S1', 'S2', 'S1', 'S2','S1', 'S2', 'S1', 'S2'],
    'enrollment_measure_name': ['M1', 'M2', 'M1', 'M2','M1', 'M2', 'M1', 'M2'],
    'value': [100, 150, 120, 170, 100, 150, 90, 100]
})

# 指定嵌套层级的列顺序
nested_result = frame_to_nested_dict_with_columns(exploded_df, ['school_code', 'district_code', 'year', 'source'])

输出结构示例

生成的字典会清晰保留每一层的列名,比如:

{
    'school_code': {
        1: {
            'district_code': {
                10: {
                    'year': {
                        2022: {
                            'source': {
                                'S1': {'M1': 100},
                                'S2': {'M2': 150}
                            }
                        },
                        2023: {
                            'source': {
                                'S1': {'M1': 120},
                                'S2': {'M2': 170}
                            }
                        }
                    }
                }
            }
        },
        2: {
            'district_code': {
                20: {
                    'year': {
                        2022: {
                            'source': {
                                'S1': {'M1': 100},
                                'S2': {'M2': 150}
                            }
                        },
                        2023: {
                            'source': {
                                'S1': {'M1': 90},
                                'S2': {'M2': 100}
                            }
                        }
                    }
                }
            }
        }
    }
}

额外优化:同时保留关联列(如school_code和school_name)

如果想在同一层级保留多个关联列(比如同时显示学校代码和名称),可以稍微修改函数支持多列分组:

def frame_to_nested_dict_with_columns(df, levels):
    if not levels:
        return df.groupby('enrollment_measure_name')['value'].first().to_dict()
    
    current_level = levels[0]
    nested_dict = {}
    
    # 处理多列作为同一层级的情况
    if isinstance(current_level, list):
        # 按多列分组,键是列值的元组
        for col_values, group in df.groupby(current_level):
            # 生成包含列名和值的标识字符串
            key_label = ", ".join([f"{col}:{val}" for col, val in zip(current_level, col_values)])
            nested_dict['school_info'] = nested_dict.get('school_info', {})
            nested_dict['school_info'][key_label] = frame_to_nested_dict_with_columns(group, levels[1:])
    else:
        # 单列层级的处理逻辑和之前一致
        for col_value, group in df.groupby(current_level):
            nested_dict[current_level] = nested_dict.get(current_level, {})
            nested_dict[current_level][col_value] = frame_to_nested_dict_with_columns(group, levels[1:])
    
    return nested_dict

调用时指定多列层级:

nested_result = frame_to_nested_dict_with_columns(exploded_df, [['school_code', 'school_name'], 'district_code', 'year', 'source'])

这样生成的字典里,school_info下会显示school_code:1, school_name:A这样的清晰标识,更便于理解。

备注:内容来源于stack exchange,提问作者nick_craft

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.16 06:54:41