如何将层级树形结构的Pandas DataFrame规范化为多列结构?
解决树形结构DataFrame的规范化问题
实现步骤
我们可以通过拆分层级标识、构建路径映射、筛选叶子节点并填充层级值这几个核心步骤完成需求:
- 拆分
Hierarchy ID为层级列表,明确每个节点的位置 - 建立层级路径与对应
Value的映射关系,方便快速取值 - 筛选出没有子节点的叶子节点(最终输出的每一行对应一条完整的叶子路径)
- 为每个叶子节点填充对应的Level_1至Level_4列值
完整代码
import pandas as pd # 示例数据(替换为你的实际数据即可) data = { 'Hierarchy ID': ['1.0', '1.1', '1.1.1', '1.1.2', '1.2', '1.2.1', '2.0', '2.1', '2.1.1', '2.1.1.1', '2.1.2'], 'Value': ['T1', 'T2', 'T3', 'T4', 'T5', 'T6', 'T7', 'T8', 'T9', 'T10', 'T11'] } df = pd.DataFrame(data) # 处理层级标识:转字符串避免数值格式问题、拆分为整数列表、计算层级深度 df['Hierarchy ID'] = df['Hierarchy ID'].astype(str) df['levels'] = df['Hierarchy ID'].str.split('.').map(lambda x: [int(i) for i in x]) df['depth'] = df['levels'].str.len() # 构建路径-值映射字典,例如(1,)对应T1,(1,1)对应T2 path_map = {} for _, row in df.iterrows(): path_tuple = tuple(row['levels']) path_map[path_tuple] = row['Value'] # 判断节点是否为叶子节点(无后续子节点) def check_leaf(row): current_path = row['levels'] current_depth = row['depth'] # 检查是否存在子节点:其他节点路径更长且前缀与当前路径一致 has_child = any(len(lvls) > current_depth and lvls[:current_depth] == current_path for lvls in df['levels']) return not has_child df['is_leaf'] = df.apply(check_leaf, axis=1) leaf_nodes = df[df['is_leaf']].copy() # 填充各Level列 max_level = df['depth'].max() for level_num in range(1, max_level + 1): col_name = f'Level_{level_num}' leaf_nodes[col_name] = leaf_nodes['levels'].map( lambda x: path_map.get(tuple(x[:level_num]), pd.NA) ) # 整理最终结果 result_df = leaf_nodes[[f'Level_{i}' for i in range(1, max_level + 1)]].reset_index(drop=True) print(result_df)
输出结果
Level_1 Level_2 Level_3 Level_4 0 T1 T2 T3 <NA> 1 T1 T2 T4 <NA> 2 T1 T5 T6 <NA> 3 T7 T8 T9 T10 4 T7 T8 T11 <NA>
内容的提问来源于stack exchange,提问作者Solijoli
相关产品推荐
相关产品推荐

