You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效获取Pandas DataFrame中评论对应的原始帖子ID?

问题描述

我有一个包含帖子(post)和评论(comment)的Pandas DataFrame,每条评论拥有id和parent id(标识其回复的帖子或评论),帖子仅包含id,数据如下:

内容类型idparent id
帖子11
评论121
评论231
评论342
评论454
帖子26
评论576

我希望为每条评论获取对应的原始帖子的id,生成包含ancestor id列的结果表:

内容类型idparent idancestor id
帖子11
评论1211
评论2311
评论3421
评论4541
帖子26
评论5766

我尝试从DataFrame末尾向前循环,迭代追溯parent id直至找到空值的parent id单元格,该方法在测试数据集上可行,但在主数据集上速度过慢。请问有更高效的实现方式吗?

原代码如下:

#creating a column for the id of the original post
df["ancestor"] = df.id
#obtaining the id of the original post for every comment
for i in reversed(range(len(df.id))): #looping trough the comments
  id = df["parent_id"][i] #variable to initialize the future loop
  parent = id
  while parent != "": #only looping trough comments
    df.ancestor[i] = id
    parent = df.parent_id[df.id == id].values[0]
    id = parent
高效实现方案

原始方案效率低的核心原因是:每次查找父节点都要遍历整个DataFrame,且重复追溯相同路径的节点时会做冗余计算。可以用字典映射+路径缓存的思路(类似并查集算法),把节点关系转换成O(1)查找的字典,同时缓存已计算的祖先结果,大幅提升处理速度。

步骤1:构建id到parent的快速映射字典

先把DataFrame的id和parent id转换成字典,避免每次查找父节点都遍历DataFrame:

# 处理空值:把空的parent id转为None,方便判断
df['parent_id'] = df['parent id'].replace('', None)
# 构建id → parent_id的映射字典
id_parent_map = df.set_index('id')['parent_id'].to_dict()

步骤2:编写带缓存的祖先查找函数

用字典缓存已经计算过的祖先结果,避免重复递归查找:

ancestor_cache = {}

def get_root_ancestor(node_id):
    parent = id_parent_map[node_id]
    # 如果当前节点是帖子(无父节点),直接标记为None
    if parent is None:
        ancestor_cache[node_id] = None
        return None
    # 如果已经缓存过该节点的祖先,直接返回
    if node_id in ancestor_cache:
        return ancestor_cache[node_id]
    # 递归查找父节点的祖先
    parent_ancestor = get_root_ancestor(parent)
    # 父节点的祖先为空,说明父节点就是原始帖子,直接用父节点id
    if parent_ancestor is None:
        ancestor_cache[node_id] = parent
    else:
        ancestor_cache[node_id] = parent_ancestor
    return ancestor_cache[node_id]

步骤3:批量生成ancestor id列

利用Pandas的apply批量处理所有节点,生成目标列:

df['ancestor id'] = df['id'].apply(get_root_ancestor)
# 可选:删除临时创建的parent_id列,恢复原数据结构
df = df.drop('parent_id', axis=1)

完整优化代码

import pandas as pd

# 示例数据
data = {
    '内容类型': ['帖子1', '评论1', '评论2', '评论3', '评论4', '帖子2', '评论5'],
    'id': [1,2,3,4,5,6,7],
    'parent id': ['', '1', '1', '2', '4', '', '6']
}
df = pd.DataFrame(data)

# 构建id到parent的映射
df['parent_id'] = df['parent id'].replace('', None)
id_parent_map = df.set_index('id')['parent_id'].to_dict()

# 缓存字典和查找函数
ancestor_cache = {}
def get_root_ancestor(node_id):
    parent = id_parent_map[node_id]
    if parent is None:
        ancestor_cache[node_id] = None
        return None
    if node_id in ancestor_cache:
        return ancestor_cache[node_id]
    parent_ancestor = get_root_ancestor(parent)
    ancestor_cache[node_id] = parent if parent_ancestor is None else parent_ancestor
    return ancestor_cache[node_id]

# 生成结果列
df['ancestor id'] = df['id'].apply(get_root_ancestor)
df = df.drop('parent_id', axis=1)

print(df)

效率提升原因

  1. 字典查找替代遍历:把父节点查找从O(n)的DataFrame遍历变成O(1)的字典查询,单次查找速度大幅提升
  2. 路径缓存避免冗余:每个节点的祖先只计算一次,后续直接复用缓存结果,避免重复递归
  3. 批量处理优化:Pandas的apply内部做了矢量化优化,比逐行循环的效率高很多

内容的提问来源于stack exchange,提问作者user18694315

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 13:45:33