如何高效获取Pandas DataFrame中评论对应的原始帖子ID?
问题描述
我有一个包含帖子(post)和评论(comment)的Pandas DataFrame,每条评论拥有id和parent id(标识其回复的帖子或评论),帖子仅包含id,数据如下:
| 内容类型 | id | parent id |
|---|---|---|
| 帖子1 | 1 | |
| 评论1 | 2 | 1 |
| 评论2 | 3 | 1 |
| 评论3 | 4 | 2 |
| 评论4 | 5 | 4 |
| 帖子2 | 6 | |
| 评论5 | 7 | 6 |
我希望为每条评论获取对应的原始帖子的id,生成包含ancestor id列的结果表:
| 内容类型 | id | parent id | ancestor id |
|---|---|---|---|
| 帖子1 | 1 | ||
| 评论1 | 2 | 1 | 1 |
| 评论2 | 3 | 1 | 1 |
| 评论3 | 4 | 2 | 1 |
| 评论4 | 5 | 4 | 1 |
| 帖子2 | 6 | ||
| 评论5 | 7 | 6 | 6 |
我尝试从DataFrame末尾向前循环,迭代追溯parent id直至找到空值的parent id单元格,该方法在测试数据集上可行,但在主数据集上速度过慢。请问有更高效的实现方式吗?
原代码如下:
#creating a column for the id of the original post df["ancestor"] = df.id #obtaining the id of the original post for every comment for i in reversed(range(len(df.id))): #looping trough the comments id = df["parent_id"][i] #variable to initialize the future loop parent = id while parent != "": #only looping trough comments df.ancestor[i] = id parent = df.parent_id[df.id == id].values[0] id = parent
高效实现方案
原始方案效率低的核心原因是:每次查找父节点都要遍历整个DataFrame,且重复追溯相同路径的节点时会做冗余计算。可以用字典映射+路径缓存的思路(类似并查集算法),把节点关系转换成O(1)查找的字典,同时缓存已计算的祖先结果,大幅提升处理速度。
步骤1:构建id到parent的快速映射字典
先把DataFrame的id和parent id转换成字典,避免每次查找父节点都遍历DataFrame:
# 处理空值:把空的parent id转为None,方便判断 df['parent_id'] = df['parent id'].replace('', None) # 构建id → parent_id的映射字典 id_parent_map = df.set_index('id')['parent_id'].to_dict()
步骤2:编写带缓存的祖先查找函数
用字典缓存已经计算过的祖先结果,避免重复递归查找:
ancestor_cache = {} def get_root_ancestor(node_id): parent = id_parent_map[node_id] # 如果当前节点是帖子(无父节点),直接标记为None if parent is None: ancestor_cache[node_id] = None return None # 如果已经缓存过该节点的祖先,直接返回 if node_id in ancestor_cache: return ancestor_cache[node_id] # 递归查找父节点的祖先 parent_ancestor = get_root_ancestor(parent) # 父节点的祖先为空,说明父节点就是原始帖子,直接用父节点id if parent_ancestor is None: ancestor_cache[node_id] = parent else: ancestor_cache[node_id] = parent_ancestor return ancestor_cache[node_id]
步骤3:批量生成ancestor id列
利用Pandas的apply批量处理所有节点,生成目标列:
df['ancestor id'] = df['id'].apply(get_root_ancestor) # 可选:删除临时创建的parent_id列,恢复原数据结构 df = df.drop('parent_id', axis=1)
完整优化代码
import pandas as pd # 示例数据 data = { '内容类型': ['帖子1', '评论1', '评论2', '评论3', '评论4', '帖子2', '评论5'], 'id': [1,2,3,4,5,6,7], 'parent id': ['', '1', '1', '2', '4', '', '6'] } df = pd.DataFrame(data) # 构建id到parent的映射 df['parent_id'] = df['parent id'].replace('', None) id_parent_map = df.set_index('id')['parent_id'].to_dict() # 缓存字典和查找函数 ancestor_cache = {} def get_root_ancestor(node_id): parent = id_parent_map[node_id] if parent is None: ancestor_cache[node_id] = None return None if node_id in ancestor_cache: return ancestor_cache[node_id] parent_ancestor = get_root_ancestor(parent) ancestor_cache[node_id] = parent if parent_ancestor is None else parent_ancestor return ancestor_cache[node_id] # 生成结果列 df['ancestor id'] = df['id'].apply(get_root_ancestor) df = df.drop('parent_id', axis=1) print(df)
效率提升原因
- 字典查找替代遍历:把父节点查找从O(n)的DataFrame遍历变成O(1)的字典查询,单次查找速度大幅提升
- 路径缓存避免冗余:每个节点的祖先只计算一次,后续直接复用缓存结果,避免重复递归
- 批量处理优化:Pandas的
apply内部做了矢量化优化,比逐行循环的效率高很多
内容的提问来源于stack exchange,提问作者user18694315
相关产品推荐
相关产品推荐

