如何基于含Total→Type→Item→Region列的DataFrame生成四级层级字典?
实现四级层级字典(Total→Type→Item→Region)
先模拟示例数据集
假设你的DataFrame结构如下,我们先构造一个样例数据方便演示:
import pandas as pd data = { 'Total': ['A', 'A', 'A', 'B', 'B'], 'Type': ['X', 'X', 'Y', 'Z', 'Z'], 'Item': ['P', 'Q', 'R', 'S', 'T'], 'Region': ['North', 'South', 'East', 'West', 'North'] } df = pd.DataFrame(data)
方法一:嵌套循环直接构建
这种方式直观,适合固定四级层级的场景:
hierarchy_dict = {} # 逐层遍历分组,构建嵌套字典 for total_val, total_group in df.groupby('Total'): hierarchy_dict[total_val] = {} for type_val, type_group in total_group.groupby('Type'): hierarchy_dict[total_val][type_val] = {} for item_val, item_group in type_group.groupby('Item'): # 提取当前Item对应的所有Region,如需去重则用unique().tolist() hierarchy_dict[total_val][type_val][item_val] = item_group['Region'].tolist()
运行后得到的字典结构如下:
{ 'A': { 'X': { 'P': ['North'], 'Q': ['South'] }, 'Y': { 'R': ['East'] } }, 'B': { 'Z': { 'S': ['West'], 'T': ['North'] } } }
方法二:递归函数构建(通用型)
如果以后需要调整层级数量,递归方式更灵活,只需修改层级参数即可:
def build_hierarchy(df, levels): # 只剩最后一层时,返回对应列的列表 if len(levels) == 1: return df[levels[0]].tolist() # 逐层分组,递归构建下一层 current_level = levels[0] next_levels = levels[1:] hierarchy = {} for key, group in df.groupby(current_level): hierarchy[key] = build_hierarchy(group, next_levels) return hierarchy # 传入层级顺序:Total→Type→Item→Region hierarchy_dict = build_hierarchy(df, ['Total', 'Type', 'Item', 'Region'])
注意事项
- 如果
Region列存在重复值,需要去重的话,把代码中的.tolist()替换为.unique().tolist() - 确保DataFrame中没有空值,否则分组时可能出现异常,必要时先做数据清洗(比如
df.dropna(subset=['Total', 'Type', 'Item', 'Region']))
内容的提问来源于stack exchange,提问作者Andreas Zaras
相关产品推荐
相关产品推荐

