如何将展开的DataFrame转换为保留列名层级的深度嵌套字典
如何将展开的DataFrame转换为保留列名层级的深度嵌套字典
我完全懂你的困扰——把扁平化的DataFrame转成嵌套字典时丢失了列名层级,确实会让字典结构变得混乱,根本没法轻松导航。别担心,我们可以修改递归逻辑,让每一层都明确保留列名作为父节点的键,而不是直接用列值当顶层键,这样整个结构就清晰多了。
解决方案:修改递归函数保留列名层级
下面是调整后的函数,它会在每一层嵌套中都保留对应的列名作为键,让你能清楚知道当前层级对应的是DataFrame的哪个字段:
import pandas as pd def frame_to_nested_dict_with_columns(df, levels): # 递归终止条件:没有剩余层级时,构建最内层的指标-值映射 if not levels: return df.groupby('enrollment_measure_name')['value'].first().to_dict() current_col = levels[0] nested_dict = {} # 按当前层级的列分组,遍历每个分组 for col_value, group in df.groupby(current_col): # 先创建以列名为键的父节点,避免覆盖 if current_col not in nested_dict: nested_dict[current_col] = {} # 递归处理剩余层级,将结果挂载到当前列值对应的位置 nested_dict[current_col][col_value] = frame_to_nested_dict_with_columns(group, levels[1:]) return nested_dict
调用示例
用你提供的exploded_df来测试,指定你想要的嵌套层级顺序(比如school_code → district_code → year → source):
# 你的原始DataFrame exploded_df = pd.DataFrame({ 'school_code': [1, 1, 1, 1, 2, 2, 2, 2], 'school_name': ['A', 'A', 'A', 'A', 'B', 'B', 'B', 'B'], 'district_code': [10, 10, 10, 10, 20, 20, 20, 20], 'year': [2022, 2022, 2023, 2023, 2022, 2022, 2023, 2023], 'source': ['S1', 'S2', 'S1', 'S2','S1', 'S2', 'S1', 'S2'], 'enrollment_measure_name': ['M1', 'M2', 'M1', 'M2','M1', 'M2', 'M1', 'M2'], 'value': [100, 150, 120, 170, 100, 150, 90, 100] }) # 指定嵌套层级的列顺序 nested_result = frame_to_nested_dict_with_columns(exploded_df, ['school_code', 'district_code', 'year', 'source'])
输出结构示例
生成的字典会清晰保留每一层的列名,比如:
{ 'school_code': { 1: { 'district_code': { 10: { 'year': { 2022: { 'source': { 'S1': {'M1': 100}, 'S2': {'M2': 150} } }, 2023: { 'source': { 'S1': {'M1': 120}, 'S2': {'M2': 170} } } } } } }, 2: { 'district_code': { 20: { 'year': { 2022: { 'source': { 'S1': {'M1': 100}, 'S2': {'M2': 150} } }, 2023: { 'source': { 'S1': {'M1': 90}, 'S2': {'M2': 100} } } } } } } } }
额外优化:同时保留关联列(如school_code和school_name)
如果想在同一层级保留多个关联列(比如同时显示学校代码和名称),可以稍微修改函数支持多列分组:
def frame_to_nested_dict_with_columns(df, levels): if not levels: return df.groupby('enrollment_measure_name')['value'].first().to_dict() current_level = levels[0] nested_dict = {} # 处理多列作为同一层级的情况 if isinstance(current_level, list): # 按多列分组,键是列值的元组 for col_values, group in df.groupby(current_level): # 生成包含列名和值的标识字符串 key_label = ", ".join([f"{col}:{val}" for col, val in zip(current_level, col_values)]) nested_dict['school_info'] = nested_dict.get('school_info', {}) nested_dict['school_info'][key_label] = frame_to_nested_dict_with_columns(group, levels[1:]) else: # 单列层级的处理逻辑和之前一致 for col_value, group in df.groupby(current_level): nested_dict[current_level] = nested_dict.get(current_level, {}) nested_dict[current_level][col_value] = frame_to_nested_dict_with_columns(group, levels[1:]) return nested_dict
调用时指定多列层级:
nested_result = frame_to_nested_dict_with_columns(exploded_df, [['school_code', 'school_name'], 'district_code', 'year', 'source'])
这样生成的字典里,school_info下会显示school_code:1, school_name:A这样的清晰标识,更便于理解。
备注:内容来源于stack exchange,提问作者nick_craft
相关产品推荐
相关产品推荐

