在Pandas DataFrame中递归计算父子比率及跟踪谱系层级
计算节点分数与父节点比率并跟踪谱系深度的解决方案
嘿,我完全get你的需求了——就是要给树形结构里的每个节点,计算它的分数和父节点分数的比率,同时还要跟踪它在整个谱系里的深度层级对吧?我给你整理了一个用Pandas实现的完整方案,刚好匹配你给出的示例数据。
首先先明确下咱们要得到的结果,基于你给的输入DataFrame,最终应该是这样的:
| id | parent_id | score | depth | ratio |
|---|---|---|---|---|
| 1 | 0 | 50 | 1 | NaN |
| 2 | 1 | 40 | 2 | 0.8 |
| 3 | 1 | 30 | 2 | 0.6 |
| 4 | 2 | 20 | 3 | 0.5 |
| 5 | 4 | 10 | 4 | 0.5 |
下面是具体的实现步骤:
1. 加载输入数据为Pandas DataFrame
首先咱们先把示例数据转换成DataFrame格式:
import pandas as pd data = { 'id': [1, 2, 3, 4, 5], 'parent_id': [0, 1, 1, 2, 4], 'score': [50, 40, 30, 20, 10] } df = pd.DataFrame(data)
2. 构建快速查找的映射字典
为了避免反复在DataFrame中查询父节点信息,咱们先创建两个字典映射,提升后续计算效率:
# 映射id到对应score,方便直接获取父节点分数 id_to_score = df.set_index('id')['score'].to_dict() # 映射id到对应parent_id,用于遍历谱系路径 id_to_parent = df.set_index('id')['parent_id'].to_dict()
3. 高效计算深度与比率(动态规划版)
如果你的数据集较大,动态规划缓存已计算的深度能避免重复遍历父节点,大幅提升效率:
# 缓存已计算的深度,id=0是根节点的虚拟父节点,深度设为0 depth_cache = {0: 0} def get_depth(node_id): # 若该节点深度已缓存,直接返回 if node_id in depth_cache: return depth_cache[node_id] # 未缓存则递归获取父节点深度,当前节点深度=父节点深度+1 parent_id = id_to_parent[node_id] depth = get_depth(parent_id) + 1 # 将计算结果存入缓存 depth_cache[node_id] = depth return depth # 为每个节点计算深度 df['depth'] = df['id'].apply(get_depth) # 计算比率:当前节点分数/父节点分数,根节点(parent_id=0)设为NaN df['ratio'] = df.apply( lambda row: row['score'] / id_to_score[row['parent_id']] if row['parent_id'] != 0 else None, axis=1 ) # 比率保留两位小数,让结果更整洁 df['ratio'] = df['ratio'].round(2)
4. 查看最终结果
运行上述代码后,打印DataFrame即可得到预期结果:
print(df)
输出:
id parent_id score depth ratio 0 1 0 50 1 NaN 1 2 1 40 2 0.8 2 3 1 30 2 0.6 3 4 2 20 3 0.5 4 5 4 10 4 0.5
这个方案逻辑清晰,适配从小规模到大规模的树形数据集。如果需要调整根节点的深度定义(比如将根节点设为0),只需修改depth_cache的初始值即可。
内容的提问来源于stack exchange,提问作者Adhi R.
相关产品推荐
相关产品推荐

