如何使用DataFrame递归查找连接及代码问题排查
递归获取DataFrame中行的完整连接关系问题解决
原始DataFrame
数据定义:
data = { 'Row_Id': [1, 2, 3, 4, 5, 6, 7], 'Inbound_Connection': [[2, 5, 3], [3], [4], [1], [], [3], [4]], 'Outbound_Connection': [[4], [1], [1, 6, 2], [3, 7], [1], [], []], 'Row_Text': ['Row 1', 'Row 2', 'Row 3', 'Row 4', 'Row 5', 'Row 6', 'Row 7'] }
表格展示:
| Row_Id | Inbound_Connection | Outbound_Connection | Row_Text |
|---|---|---|---|
| 1 | [2, 5, 3] | [4] | Row 1 |
| 2 | [3] | [1] | Row 2 |
| 3 | [4] | [1, 6, 2] | Row 3 |
| 4 | [1] | [3, 7] | Row 4 |
| 5 | [] | [1] | Row 5 |
| 6 | [3] | [] | Row 6 |
| 7 | [4] | [] | Row 7 |
需求
给定任意行ID,递归获取该行的Inbound_Connection和Outbound_Connection,并继续递归获取这些连接节点的入站、出站连接,形成完整的层级关系。
现有代码问题
当前代码中,当递归到已访问过的节点时,直接返回空字典{},导致该节点的基础信息(如Row_Text)以及其连接关系被省略,从而出现多元素连接或嵌套连接时信息缺失的情况。
修正方案
修改递归逻辑:当节点已被访问时,返回该节点的基础信息(仅Row_Text),而非空字典;同时保留visited集合避免循环递归,但不丢失节点本身的信息。
修正后的代码:
import pandas as pd import json data = { 'Row_Id': [1, 2, 3, 4, 5, 6, 7], 'Inbound_Connection': [[2, 5, 3], [3], [4], [1], [], [3], [4]], 'Outbound_Connection': [[4], [1], [1, 6, 2], [3, 7], [1], [], []], 'Row_Text': ['Row 1', 'Row 2', 'Row 3', 'Row 4', 'Row 5', 'Row 6', 'Row 7'] } df = pd.DataFrame(data) def find_connections(df, row_id, visited=None): if visited is None: visited = set() # 获取当前行基础信息 row = df[df['Row_Id'] == row_id] if row.empty: return {} row_text = row['Row_Text'].values[0] # 已访问节点仅返回基础信息,避免循环递归 if row_id in visited: return {'Row_Text': row_text} visited.add(row_id) inbound_connections = row['Inbound_Connection'].values[0] outbound_connections = row['Outbound_Connection'].values[0] connections_dict = {'Row_Text': row_text} # 处理入站连接 if inbound_connections: connections_dict['Inbound'] = {} for conn_id in inbound_connections: connections_dict['Inbound'][conn_id] = find_connections(df, conn_id, visited) # 处理出站连接 if outbound_connections: connections_dict['Outbound'] = {} for conn_id in outbound_connections: connections_dict['Outbound'][conn_id] = find_connections(df, conn_id, visited) return connections_dict # 生成所有行的完整连接关系 connections_dict = {row_id: find_connections(df, row_id) for row_id in df['Row_Id']} # 打印格式化后的结果 print(json.dumps(connections_dict, indent=2))
修正后效果
所有连接节点的Row_Text信息均被保留,嵌套连接的层级关系完整呈现,不会出现空字典导致的信息缺失。以Row_Id=1的输出片段为例:
{ "1": { "Row_Text": "Row 1", "Inbound": { "2": { "Row_Text": "Row 2", "Inbound": { "3": { "Row_Text": "Row 3", "Inbound": { "4": { "Row_Text": "Row 4", "Inbound": { "1": { "Row_Text": "Row 1" } }, "Outbound": { "3": { "Row_Text": "Row 3" }, "7": { "Row_Text": "Row 7", "Inbound": { "4": { "Row_Text": "Row 4" } } } } } }, "Outbound": { "1": { "Row_Text": "Row 1" }, "6": { "Row_Text": "Row 6", "Inbound": { "3": { "Row_Text": "Row 3" } } }, "2": { "Row_Text": "Row 2" } } } }, "Outbound": { "1": { "Row_Text": "Row 1" } } }, "5": { "Row_Text": "Row 5", "Outbound": { "1": { "Row_Text": "Row 1" } } }, "3": { "Row_Text": "Row 3" } }, "Outbound": { "4": { "Row_Text": "Row 4" } } } }
内容的提问来源于stack exchange,提问作者Lav Sharma
相关产品推荐
相关产品推荐

