如何在Pandas&Graphviz家谱生成程序中实现配偶节点相邻
解决Graphviz生成家谱时配偶节点不相邻的问题
数据结构(CSV)
| ID | S | First name | Last name | DoB | DoD | FatherID | MotherID | SpouseID | Place of birth | Job |
|---|---|---|---|---|---|---|---|---|---|---|
| JoS1 | M | John | S | 1111 | 2222 | MaS1 | India | Job-1 | ||
| MaS1 | F | Mary | S | 1112 | JoS1 | India | Job-2 | |||
| JaS | M | Jacob | S | 1113 | JoS1 | MaS1 | KeS | India | Job-3 | |
| JoS2 | M | Joe | S | 1114 | 2225 | JoS1 | MaS1 | AnS | India | Job-4 |
| MaS2 | F | Macy | D | 1115 | JoS1 | MaS1 | AnD | India | Job-5 | |
| KeS | F | Keysha | S | 1116 | JaS | India | Job-6 | |||
| AnD | M | Andy | D | 1117 | MaS2 | India | Job-7 | |||
| AnS | F | Anna | S | 1118 | JoS2 | India | Job-8 | |||
| MiS | M | Mike | S | 1119 | JaS | KeS | India | |||
| SaS | M | Sam | S | 1120 | JaS | KeS | India | |||
| MaS3 | F | Matt | S | 2345 | JoS2 | AnS | India |
原代码
from graphviz import Digraph import pandas as pd import numpy as np rawdf = pd.read_csv('/content/drive/MyDrive/ftdata.csv', keep_default_na=False) ## Change file path el1 = rawdf[['ID','MotherID','SpouseID']] el2 = rawdf[['ID','FatherID','SpouseID']] el1.columns = ['Child', 'ParentID','SpouseID'] el2.columns = el1.columns el = pd.concat([el1, el2]) el.replace('', np.nan, regex=True, inplace = True) t = pd.DataFrame({'tmp':['no_entry'+str(i) for i in range(el.shape[0])]}) el['ParentID'].fillna(t['tmp'], inplace=True) el['SpouseID'].fillna(t['tmp'], inplace=True) df = el.merge(rawdf, left_index=True, right_index=True, how='left') df['name'] = df[df.columns[4:6]].apply(lambda x: ' '.join(x.dropna().astype(str)),axis=1) df = df.drop(['Child','FatherID', 'ID', 'First name', 'Last name'], axis=1) df = df[['ID', 'name', 'S', 'DoB', 'DoD', 'Place of birth', 'Job', 'ParentID']] f = Digraph('neato', format='jpg', encoding='utf8', filename='testfile', node_attr={'style': 'filled'}, graph_attr={"concentrate": "true", "splines":"ortho"}) f.attr('node', shape='box') for index, row in df.iterrows(): f.node(row['ID'], label= str(row['name']) + '\n' + str(row['Job']) + '\n'+ str(row['DoB']) + '\n' + str(row['Place of birth']) + '\n†' + str(row['DoD']), _attributes={'color':'lightpink' if row['S']=='F' else 'lightblue'if row['S']=='M' else 'lightgray'}) for index, row in df.iterrows(): f.edge(str(row["ParentID"]), str(row["ID"]), label='') f.view()
问题描述
当前生成的家谱中,配偶节点未被分组相邻,导致家谱结构可读性差,需要调整布局让配偶节点显示在同一水平位置且相邻。
解决方案
核心修改点
- 去重节点:原代码因合并父母系数据导致节点重复,需去重避免重复渲染
- 识别唯一配偶对:避免重复处理双向配偶关系(如A→B和B→A只处理一次)
- 子图强制同层级:通过
subgraph给配偶节点添加rank=same属性,强制处于同一水平行 - 区分配偶与亲子边:用无方向虚线连接配偶,和亲子实线做区分
- 切换布局引擎:改用
dot层级布局,比neato力导向布局更适配家谱结构
修改后的完整代码
from graphviz import Digraph import pandas as pd import numpy as np rawdf = pd.read_csv('/content/drive/MyDrive/ftdata.csv', keep_default_na=False) ## Change file path # 处理亲子关系数据 el1 = rawdf[['ID','MotherID','SpouseID']] el2 = rawdf[['ID','FatherID','SpouseID']] el1.columns = ['Child', 'ParentID','SpouseID'] el2.columns = el1.columns el = pd.concat([el1, el2]) el.replace('', np.nan, regex=True, inplace = True) t = pd.DataFrame({'tmp':['no_entry'+str(i) for i in range(el.shape[0])]}) el['ParentID'].fillna(t['tmp'], inplace=True) el['SpouseID'].fillna(t['tmp'], inplace=True) df = el.merge(rawdf, left_index=True, right_index=True, how='left') df['name'] = df[['First name', 'Last name']].apply(lambda x: ' '.join(x.dropna().astype(str)),axis=1) df = df.drop(['Child','FatherID', 'ID_y', 'First name', 'Last name'], axis=1) df = df.rename(columns={'ID_x': 'ID'}) df = df[['ID', 'name', 'S', 'DoB', 'DoD', 'Place of birth', 'Job', 'ParentID', 'SpouseID']] # 去重,避免同一节点重复处理 df = df.drop_duplicates(subset=['ID']).reset_index(drop=True) # 初始化Graphviz,改用dot层级布局 f = Digraph('dot', format='jpg', encoding='utf8', filename='family_tree', node_attr={'style': 'filled'}, graph_attr={"concentrate": "true", "splines":"ortho", "rankdir": "TB"}) f.attr('node', shape='box') # 添加所有节点,优化标签显示 for index, row in df.iterrows(): dod_text = f'\n†{row["DoD"]}' if row["DoD"] else '' label = (f'{row["name"]}\n{row["Job"]}\n{row["DoB"]}\n{row["Place of birth"]}' f'{dod_text}') f.node(row['ID'], label=label, color='lightpink' if row['S']=='F' else 'lightblue' if row['S']=='M' else 'lightgray') # 添加亲子边,跳过空父/母ID for index, row in df.iterrows(): if row['ParentID'] not in t['tmp'].values: f.edge(str(row["ParentID"]), str(row["ID"]), label='') # 处理配偶关系,创建唯一配偶对 processed_couples = set() for index, row in df.iterrows(): spouse_id = row['SpouseID'] if spouse_id not in t['tmp'].values and (spouse_id, row['ID']) not in processed_couples: # 子图设置配偶同层级 with f.subgraph() as s: s.attr(rank='same') s.node(row['ID']) s.node(spouse_id) # 添加无方向虚线连接配偶 f.edge(row['ID'], spouse_id, style='dashed', dir='none') processed_couples.add((row['ID'], spouse_id)) f.view()
修改效果说明
- 配偶节点会被强制放在同一水平行,相邻显示
- 用虚线区分配偶关系,实线区分亲子关系,结构更清晰
- 优化了空死亡日期的标签显示,避免出现无效字符
dot布局让家谱保持垂直层级结构,符合传统家谱阅读习惯
内容的提问来源于stack exchange,提问作者Kshitij Khandelwal
相关产品推荐
相关产品推荐

