使用NetworkX扁平化父子层级并添加额外列的技术问题
为NetworkX扁平化的父子层级结果添加Description列
我已经用NetworkX实现了父子层级的扁平化,输出了各层级节点的DataFrame,但不知道如何将原数据中的Description列对应添加到结果中。现有代码如下:
import pandas as pd data = [ ['Unit A', 'Department Q', 'Lorem'], ['Unit A', 'Department R', 'Ipsum'], ['Unit A', 'Department S', 'dolor'], ['Department S', 'Office 1', 'sit'], ['Department S', 'Office 2', 'amet'], ['Unit B', 'Department X', 'consetetur'], ['Unit B', 'Department Y', 'sadipscing'], ['Unit B', 'Department Z', 'elitr'], ['Department Z', 'Office 3', 'sed'], ['Department Z', 'Office 4', 'diam'], ['Office 4', 'Place K', 'nonumy'] , ['Office 4', 'Place L', 'eirmod'] ] df = pd.DataFrame(data, columns=['Parent', 'Child', 'Description'])
接着创建层级结构并转换为DataFrame:
import networkx as nx G = nx.from_pandas_edgelist(df, source='Parent', target='Child', create_using=nx.DiGraph, edge_attr=True, ) roots = (v for v, d in G.in_degree() if d == 0) leaves = [v for v, d in G.out_degree() if d == 0] out = (pd.DataFrame(path for root in roots for path in nx.all_simple_paths(G, root, leaves)) .add_prefix('Node_') ) print(out)
当前输出结果:
Node_0 Node_1 Node_2 Node_3 0 Unit A Department Q None None 1 Unit A Department R None None 2 Unit A Department S Office 1 None 3 Unit A Department S Office 2 None 4 Unit B Department X None None 5 Unit B Department Y None None 6 Unit B Department Z Office 3 None 7 Unit B Department Z Office 4 Place K 8 Unit B Department Z Office 4 Place L
解决方案
原数据中的Description是每条父-子边的属性,我们可以在生成路径时,同时提取每条边对应的描述,再将这些描述作为额外列添加到结果DataFrame中。
修改后的完整代码:
import pandas as pd import networkx as nx data = [ ['Unit A', 'Department Q', 'Lorem'], ['Unit A', 'Department R', 'Ipsum'], ['Unit A', 'Department S', 'dolor'], ['Department S', 'Office 1', 'sit'], ['Department S', 'Office 2', 'amet'], ['Unit B', 'Department X', 'consetetur'], ['Unit B', 'Department Y', 'sadipscing'], ['Unit B', 'Department Z', 'elitr'], ['Department Z', 'Office 3', 'sed'], ['Department Z', 'Office 4', 'diam'], ['Office 4', 'Place K', 'nonumy'] , ['Office 4', 'Place L', 'eirmod'] ] df = pd.DataFrame(data, columns=['Parent', 'Child', 'Description']) G = nx.from_pandas_edgelist(df, source='Parent', target='Child', create_using=nx.DiGraph, edge_attr=True, ) roots = (v for v, d in G.in_degree() if d == 0) leaves = [v for v, d in G.out_degree() if d == 0] # 生成包含节点和对应边描述的字典列表 path_data = [] for root in roots: for path in nx.all_simple_paths(G, root, leaves): # 先构建节点列的键值对 row = {'Node_' + str(i): node for i, node in enumerate(path)} # 提取每条边的Description,存入对应的Desc列 for i in range(len(path)-1): parent_node = path[i] child_node = path[i+1] row[f'Desc_{i}'] = G[parent_node][child_node]['Description'] path_data.append(row) # 转换为DataFrame,缺失的列自动填充为None out = pd.DataFrame(path_data) print(out)
最终输出结果
Node_0 Node_1 Node_2 Node_3 Desc_0 Desc_1 Desc_2 0 Unit A Department Q None None Lorem None None 1 Unit A Department R None None Ipsum None None 2 Unit A Department S Office 1 None dolor sit None 3 Unit A Department S Office 2 None dolor amet None 4 Unit B Department X None None consetetur None None 5 Unit B Department Y None None sadipscing None None 6 Unit B Department Z Office 3 None elitr sed None 7 Unit B Department Z Office 4 Place K elitr diam nonumy 8 Unit B Department Z Office 4 Place L elitr diam eirmod
关键说明
- 遍历每条路径时,先构建节点的键值对(如
Node_0、Node_1) - 对路径中每一对相邻节点,从图中提取对应的
Description属性,存入Desc_0、Desc_1等列 - 用
pd.DataFrame()转换时,缺失的列会自动填充为None,保证结果格式统一
内容的提问来源于stack exchange,提问作者Frank
相关产品推荐
相关产品推荐

