如何用Pandas生成Python graphviz Digraph所需指定格式依赖数据
Pandas流程依赖数据转换适配Graphviz绘图方案
现有逻辑
已通过Pandas分组聚合得到流程依赖映射字典,对应代码及输出如下:
dependent = df.groupby('dependent_processid')['Processid'].apply(list).to_dict() print(dependent) #output: {6720: [6721], 6721: [6722, 6724, 6725], 6725: [6723, 6726], 6753: [7177]}
目标格式
需要输出父流程ID -> {关联子流程ID集合}格式的依赖关系,供graphviz.Digraph绘制树形流程图使用,目标输出示例:
6720-> {6721} 6721-> {6722, 6724, 6725} 6725-> {6723, 6726} ....and so on...
实现方法
1. 直接输出指定格式文本
遍历已生成的依赖字典,将子流程列表转为集合后按要求格式化打印即可,无需额外Pandas处理:
for parent_pid, child_pids in dependent.items(): print(f"{parent_pid}-> {set(child_pids)}")
执行后输出完全匹配目标格式:
6720-> {6721} 6721-> {6724, 6722, 6725} 6725-> {6723, 6726} 6753-> {7177}
注:集合是无序结构,输出时子ID的顺序可能和原列表不一致,不影响绘图逻辑。如果需要保留原有顺序,把
set(child_pids)替换为child_pids即可。
2. 直接对接Digraph绘图(无需转文本格式)
实际绘图时不需要特意输出上述文本格式,直接遍历依赖字典批量添加节点和有向边即可,代码更简洁:
from graphviz import Digraph # 初始化有向图 flow_dot = Digraph(name='process_dag', node_attr={'shape': 'box'}) # 批量添加边和节点 for parent, children in dependent.items(): flow_dot.node(str(parent)) for child in children: flow_dot.node(str(child)) flow_dot.edge(str(parent), str(child)) # 生成并查看流程图 flow_dot.render(view=True)
注意:graphviz节点ID默认接收字符串类型,传入整数时主动转str可避免类型兼容问题。
内容的提问来源于stack exchange,提问作者Akshay Lokhande
相关产品推荐
相关产品推荐

