如何从Pandas DataFrame生成多层级字典?
构建多层级字典实现方案
背景
给定以下输入数据和生成的DataFrame:
my_list = [['Japan', 'Consumer Durables'], ['United States', 'Electronic Technology', 'ANALYST PICK'], ['Japan', 'Finance'], ['South Korea', 'Electronic Technology']] df = pd.DataFrame(my_list, columns=['country','sector','flag'])
对应的DataFrame结构:
country sector flag 0 Japan Consumer Durables None 1 United States Electronic Technology ANALYST PICK 2 Japan Finance None 3 South Korea Electronic Technology None
需要生成一个多层级字典,其中indices对应原DataFrame的行索引,期望输出结构如下:
{"groups" : [ { "name" : "Japan", "groups" : [ {"name" : "Consumer Durables", "indices" : [0]}, {"name" : "Finance", "indices" : [2]} ], }, { "name" : "United States", "groups" : [ { "name" : "Electronic Technology", "groups" : [ {"name" : "ANALYST PICK", "indices" : [1]} ] } ] }, { "name" : "South Korea", "groups" : [ {"name" : "Electronic Technology", "indices" : [3]} ] } ] }
解决方案
可以通过分层遍历+节点匹配的方式实现,核心是按country→sector→flag的层级处理,忽略None值的节点,同时收集对应索引。
实现代码
import pandas as pd my_list = [['Japan', 'Consumer Durables'], ['United States', 'Electronic Technology', 'ANALYST PICK'], ['Japan', 'Finance'], ['South Korea', 'Electronic Technology']] df = pd.DataFrame(my_list, columns=['country','sector','flag']) def build_hierarchy(df, level_order): root = {"groups": []} for idx, row in df.iterrows(): # 生成当前行的有效层级路径,跳过None值 path = [row[level] for level in level_order if row[level] is not None] current_groups = root["groups"] for i, node_name in enumerate(path): # 检查当前层级是否已有该节点 match = next((n for n in current_groups if n["name"] == node_name), None) if not match: # 创建新节点:最后一层带indices,其他层带groups if i == len(path) - 1: new_node = {"name": node_name, "indices": [idx]} else: new_node = {"name": node_name, "groups": []} current_groups.append(new_node) # 更新当前处理的层级组(非最后一层才继续) current_groups = new_node.get("groups", []) if i != len(path)-1 else [] else: # 节点已存在:最后一层追加索引,否则进入子层级 if i == len(path) - 1: match["indices"].append(idx) current_groups = match.get("groups", []) return root # 指定层级顺序:country -> sector -> flag output = build_hierarchy(df, ["country", "sector", "flag"]) print(output)
代码说明
- 路径生成:遍历每一行时,先过滤掉值为
None的列,得到当前行的有效层级路径(比如美国行的路径是["United States", "Electronic Technology", "ANALYST PICK"])。 - 节点匹配与创建:沿着路径逐层查找节点:
- 若节点不存在,根据是否为最后一层创建对应结构(最后一层包含
indices列表,其他层包含子groups)。 - 若节点已存在,最后一层则追加当前索引,否则进入该节点的子
groups继续处理。
- 若节点不存在,根据是否为最后一层创建对应结构(最后一层包含
- 层级跟踪:通过
current_groups变量跟踪当前处理的层级列表,实现嵌套结构的构建。
运行代码后即可得到符合要求的多层级字典。
内容的提问来源于stack exchange,提问作者Sungmin Son
相关产品推荐
相关产品推荐

