You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Pandas DataFrame递归映射为指定结构的Python字典?

递归将Pandas DataFrame转换为层级嵌套字典

需求描述

需要将给定的Pandas DataFrame递归映射为特定层级结构的Python字典:

  • 每个层级包含details列表,汇总对应分组的total_hc(hc列求和)和response_total(response列求和)
  • 子层级嵌套在当前层级的rows字段中

输入DataFrame

import pandas as pd

data = {
  "group1": ["A", "A", "B", "B"],
  "group2": ["grp1", "grp2", "grp1", "grp2"],
  "hc": [50, 40, 45, 90],
  "response": [12, 30, 43, 80]
}

df = pd.DataFrame(data)

期望输出结构

output = {
    "rows":[
        {
        "details": [{        
            "level": "A",
            "total_hc": 90,
            "response_total": 42
        }],
            "rows":[
                {
                "details": [{        
                    "level": "grp1",
                    "total_hc": 50,
                    "response_total": 12
                }]
                },
                {
                "details": [{        
                    "level": "grp2",
                    "total_hc": 40,
                    "response_total": 30
                }]
                }
            ]
        },
        {
        "details": [{        
            "level": "B",
            "total_hc": 135,
            "response_total": 123
        }],
            "rows":[
                {
                "details": [{        
                    "level": "grp1",
                    "total_hc": 45,
                    "response_total": 43
                }]
                },
                {
                "details": [{        
                    "level": "grp2",
                    "total_hc": 90,
                    "response_total": 80
                }]
                }
            ]
        }
    ]
}

已尝试的分组代码

group_df = df.groupby(["group1", "group2"]).sum()
group_df.to_dict("index")

解决方案

可以通过递归函数实现层级结构的构建,核心思路是按分组层级逐步向下处理,每一层生成对应的details汇总,再递归处理子分组生成rows:

def build_hierarchy(df, group_columns):
    # 无剩余分组列时返回None,代表当前层级无子rows
    if not group_columns:
        return None
    
    current_group = group_columns[0]
    remaining_groups = group_columns[1:]
    
    hierarchy_rows = []
    # 按当前分组列遍历分组结果
    for group_val, group_data in df.groupby(current_group):
        # 计算当前分组的汇总值
        total_hc = group_data['hc'].sum()
        response_total = group_data['response'].sum()
        
        # 构建当前层级的details
        detail = {
            "level": group_val,
            "total_hc": total_hc,
            "response_total": response_total
        }
        
        # 递归处理剩余分组,生成子层级rows
        child_rows = build_hierarchy(group_data, remaining_groups)
        
        # 组装当前层级的字典结构
        current_level = {"details": [detail]}
        if child_rows is not None:
            current_level["rows"] = child_rows
        
        hierarchy_rows.append(current_level)
    
    return hierarchy_rows

# 调用函数,指定分组层级顺序:group1 -> group2
result = {"rows": build_hierarchy(df, ["group1", "group2"])}

# 格式化输出验证结果
import json
print(json.dumps(result, indent=2))

代码说明

  1. build_hierarchy函数接收DataFrame和分组列列表,按顺序逐层处理分组
  2. 每一层先计算当前分组的hc和response总和,生成details内容
  3. 递归调用函数处理剩余的分组列,生成子层级的rows
  4. 最终将最顶层的rows包装成目标字典结构

内容的提问来源于stack exchange,提问作者Joanna

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 10:52:38