You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas DataFrame重复ID处理:合并或替换含新Comments的行

解决方案:处理Pandas中同ID帖子的评论更新问题

针对你遇到的同ID帖子评论更新时无法正确替换或合并的问题,这里提供两种实用方案:

方案1:直接替换旧行(保留最新的帖子数据)

如果新获取的new_posts中,同ID的帖子包含了最新的评论及其他字段内容,直接保留最后出现的行即可覆盖旧数据:

posts = pd.DataFrame(data, columns=['ID', 'Date/Time', 'Title', 'Body', 'Comments'])
new_posts = get_new_posts()

# 合并后按ID去重,保留最后一行(即新获取的帖子数据)
combined_posts = pd.concat([posts, new_posts], ignore_index=True)
posts = combined_posts.drop_duplicates('ID', keep='last')

原理:drop_duplicates的keep='last'参数会保留每个ID最后一次出现的行,由于new_posts在concat时放在后面,新数据会自动覆盖旧数据。

方案2:合并同ID的评论列表(保留旧评论,新增新评论)

如果只需要合并评论,其他字段(标题、内容等)保持不变,可以通过分组聚合实现:

posts = pd.DataFrame(data, columns=['ID', 'Date/Time', 'Title', 'Body', 'Comments'])
new_posts = get_new_posts()

combined_posts = pd.concat([posts, new_posts], ignore_index=True)

# 按ID分组,合并评论列表(去重),其他字段取首次/末次值(根据需求调整)
posts = combined_posts.groupby('ID', as_index=False).agg(
    **{'Date/Time': ('Date/Time', 'last'),  # 取最新的时间
       'Title': ('Title', 'first'),        # 标题不变取首次
       'Body': ('Body', 'first'),          # 内容不变取首次
       'Comments': ('Comments', lambda x: list(set().union(*x)))}  # 合并所有评论并去重
)

注意事项:

  • 如果评论不需要去重,把lambda x: list(set().union(*x))改成lambda x: sum(x, [])即可,但会保留重复的评论ID。
  • 其他字段的聚合方式可根据实际需求调整,比如'last'取最新值,'first'取原始值。

内容的提问来源于stack exchange,提问作者bemy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 11:17:12