You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Pandas DataFrame生成多层级字典?

构建多层级字典实现方案

背景

给定以下输入数据和生成的DataFrame:

my_list = [['Japan', 'Consumer Durables'], ['United States', 'Electronic Technology', 'ANALYST PICK'], ['Japan', 'Finance'], ['South Korea', 'Electronic Technology']]
df = pd.DataFrame(my_list, columns=['country','sector','flag'])

对应的DataFrame结构:

country               sector           flag
0      Japan   Consumer Durables          None
1  United States  Electronic Technology  ANALYST PICK
2      Japan               Finance          None
3  South Korea   Electronic Technology          None

需要生成一个多层级字典,其中indices对应原DataFrame的行索引,期望输出结构如下:

{"groups" :
    [
        {
            "name" : "Japan",
            "groups" :
                [
                    {"name" : "Consumer Durables", "indices" : [0]},
                    {"name" : "Finance", "indices" : [2]}
                ],
        },
        {
            "name" : "United States",
            "groups" :
                [
                    {
                        "name" : "Electronic Technology",
                        "groups" :
                            [
                                {"name" : "ANALYST PICK", "indices" : [1]}
                            ]
                    }
                ]
        },
        {
            "name" : "South Korea",
            "groups" :
                [
                    {"name" : "Electronic Technology", "indices" : [3]}
                ]
        }
    ]
}

解决方案

可以通过分层遍历+节点匹配的方式实现,核心是按country→sector→flag的层级处理,忽略None值的节点,同时收集对应索引。

实现代码

import pandas as pd

my_list = [['Japan', 'Consumer Durables'], ['United States', 'Electronic Technology', 'ANALYST PICK'], ['Japan', 'Finance'], ['South Korea', 'Electronic Technology']]
df = pd.DataFrame(my_list, columns=['country','sector','flag'])

def build_hierarchy(df, level_order):
    root = {"groups": []}
    
    for idx, row in df.iterrows():
        # 生成当前行的有效层级路径,跳过None值
        path = [row[level] for level in level_order if row[level] is not None]
        current_groups = root["groups"]
        
        for i, node_name in enumerate(path):
            # 检查当前层级是否已有该节点
            match = next((n for n in current_groups if n["name"] == node_name), None)
            
            if not match:
                # 创建新节点:最后一层带indices,其他层带groups
                if i == len(path) - 1:
                    new_node = {"name": node_name, "indices": [idx]}
                else:
                    new_node = {"name": node_name, "groups": []}
                current_groups.append(new_node)
                # 更新当前处理的层级组(非最后一层才继续)
                current_groups = new_node.get("groups", []) if i != len(path)-1 else []
            else:
                # 节点已存在:最后一层追加索引,否则进入子层级
                if i == len(path) - 1:
                    match["indices"].append(idx)
                current_groups = match.get("groups", [])
    
    return root

# 指定层级顺序:country -> sector -> flag
output = build_hierarchy(df, ["country", "sector", "flag"])
print(output)

代码说明

  1. 路径生成:遍历每一行时,先过滤掉值为None的列,得到当前行的有效层级路径(比如美国行的路径是["United States", "Electronic Technology", "ANALYST PICK"])。
  2. 节点匹配与创建:沿着路径逐层查找节点:
    • 若节点不存在,根据是否为最后一层创建对应结构(最后一层包含indices列表,其他层包含子groups)。
    • 若节点已存在,最后一层则追加当前索引,否则进入该节点的子groups继续处理。
  3. 层级跟踪:通过current_groups变量跟踪当前处理的层级列表,实现嵌套结构的构建。

运行代码后即可得到符合要求的多层级字典。


内容的提问来源于stack exchange,提问作者Sungmin Son

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 07:55:23