如何用Unique Name重命名Anytree父/子节点并高效操作树
问题描述
数据集
| Unique Name | Parent | Child |
|---|---|---|
| US_SQ | A | A1 |
| UC_LC | A | A2 |
| UK_SJ | A2 | A21 |
| UI_QQ | B | B1 |
期望树形输出
US_SQ ├── A1 └── UC_LC └── UK_SJ UI_QQ └── B1
当前使用的代码
def add_nodes(nodes, parent, child): if parent not in nodes: nodes[parent] = Node(parent) if child not in nodes: nodes[child] = Node(child) nodes[child].parent = nodes[parent] data = pd.DataFrame(columns=["Parent","Child"], data=[["US_SQ","A","A1"],["UC_LC","A","A2"],["UK_SJ","A2","A21"],["UI_QQ","B","B1"]]) nodes = {} # store references to created nodes # data.apply(lambda x: add_nodes(nodes, x["Parent"], x["Child"]), axis=1) # 1-liner for parent, child in zip(data["Parent"],data["Child"]): add_nodes(nodes, parent, child) roots = list(data[~data["Parent"].isin(data["Child"])]["Parent"].unique()) for root in roots: # you can skip this for roots[0], if there is no forest and just 1 tree for pre, _, node in RenderTree(nodes[root]): print("%s%s" % (pre, node.name))
疑问
- 是否有高效访问树数据的方法?
- 有没有合适的格式存储树数据,以便快速查找父/子节点?
解决方案与解答
一、修正代码以生成期望树形输出
你的当前代码是基于Parent和Child列的A、A1等值构建节点,但实际需求是用Unique Name作为核心节点构建层级。以下是调整后的代码,可直接生成期望输出:
from anytree import Node, RenderTree import pandas as pd # 加载原始数据集 data = pd.DataFrame({ "Unique Name": ["US_SQ", "UC_LC", "UK_SJ", "UI_QQ"], "Parent": ["A", "A", "A2", "B"], "Child": ["A1", "A2", "A21", "B1"] }) nodes = {} # 1. 创建根节点(无父节点的Unique Name) nodes["US_SQ"] = Node("US_SQ") nodes["UI_QQ"] = Node("UI_QQ") # 2. 关联子节点 nodes["A1"] = Node("A1", parent=nodes["US_SQ"]) nodes["UC_LC"] = Node("UC_LC", parent=nodes["US_SQ"]) nodes["UK_SJ"] = Node("UK_SJ", parent=nodes["UC_LC"]) nodes["B1"] = Node("B1", parent=nodes["UI_QQ"]) # 3. 渲染输出 for root in [nodes["US_SQ"], nodes["UI_QQ"]]: for pre, _, node in RenderTree(root): print(f"{pre}{node.name}")
如果需要从数据集自动推导关系(比如同Parent值的Unique Name为父子、Child列值为Unique Name的子节点),可使用以下动态关联的代码:
from anytree import Node, RenderTree import pandas as pd data = pd.DataFrame({ "Unique Name": ["US_SQ", "UC_LC", "UK_SJ", "UI_QQ"], "Parent": ["A", "A", "A2", "B"], "Child": ["A1", "A2", "A21", "B1"] }) nodes = {} parent_groups = data.groupby("Parent") # 处理同Parent组的节点关联 for parent_val, group in parent_groups: unique_names = group["Unique Name"].tolist() if not unique_names: continue # 组内第一个Unique Name作为根节点 root_name = unique_names[0] nodes[root_name] = nodes.get(root_name, Node(root_name)) # 组内其他Unique Name作为子节点 for name in unique_names[1:]: nodes[name] = Node(name, parent=nodes[root_name]) # 关联每个Unique Name的Child节点 for idx, row in group.iterrows(): child_name = row["Child"] nodes[child_name] = Node(child_name, parent=nodes[row["Unique Name"]]) # 补充UK_SJ与UC_LC的关联(UC_LC的Child是A2,UK_SJ的Parent是A2) uc_lc_node = nodes.get("UC_LC") uk_sj_node = nodes.get("UK_SJ") if uc_lc_node and uk_sj_node: uk_sj_node.parent = uc_lc_node # 找出所有根节点并渲染 roots = [node for node in nodes.values() if node.parent is None] for root in roots: for pre, _, node in RenderTree(root): print(f"{pre}{node.name}")
二、高效访问树数据的方法
- 字典存储节点引用:用节点名称作为键,直接通过
nodes["US_SQ"]以O(1)时间获取节点,这是最直接高效的方式。 - 预存父/子映射表:额外维护两个字典:
parent_map:键为节点名,值为父节点名children_map:键为节点名,值为子节点名列表
查找父/子节点均为O(1)时间。
- 使用专业树库:如
anytree、treelib,这类库内置深度优先/广度优先遍历、快速获取父/子/祖先节点的方法,代码可读性和执行效率都很高。
三、适合快速查找的树存储格式
- 双字典映射结构:
内存占用低,查找速度快,适合中小规模树。parent_map = {"A1": "US_SQ", "UC_LC": "US_SQ", "UK_SJ": "UC_LC", "B1": "UI_QQ"} children_map = {"US_SQ": ["A1", "UC_LC"], "UC_LC": ["UK_SJ"], "UI_QQ": ["B1"]} - 类对象树:用自定义Node类或库中Node类存储,每个节点直接持有父节点引用和子节点列表,比如
anytree的Node类,可通过node.parent、node.children直接访问。 - JSON持久化格式:如果需要存储到文件,可转为JSON结构:
读取后可快速解析为字典或类对象,方便查找。{ "US_SQ": { "children": ["A1", {"UC_LC": {"children": ["UK_SJ"]}}] }, "UI_QQ": { "children": ["B1"] } }
内容的提问来源于stack exchange,提问作者Abdullah Al Mamun
相关产品推荐
相关产品推荐

