You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按推文会话关联规则对Pandas DataFrame或Python列表重排序

推文会话关联排序实现

需求背景

现有按创建时间升序排列的推文数据,包含两个字段:

  • reference:被引用/回复的推文ID
  • uid:当前推文自身ID
    如果reference等于uid,说明是会话起始节点;否则是对其他推文的回复。需要排序后同时满足:
  1. 整体保留推文的发布时间先后逻辑
  2. 同一会话的所有回复、子回复聚合排列,紧跟在被回复推文的后方

原始数据结构

import pandas as pd

data = [
 (638009197035522, 655784141500417), # 0
 (693075572527105, 693075572527105), # 1
 (655784141500417, 693668642918400), # 2
 (693075572527105, 694397537353729), # 3
 (694397537353729, 695737600794624), # 4
 (695737600794624, 700168400654337), # 5
 (693075572527105, 929811762360322), # 6
 (929811762360322, 931830115979265), # 7
 (931830115979265, 951912745500672), # 8
 (951912745500672, 965073687117824)] # 9

df = pd.DataFrame(data, columns=['reference', 'uid'])

预期输出

[(638009197035522, 655784141500417),
 (655784141500417, 693668642918400),
 (693075572527105, 693075572527105),
 (693075572527105, 694397537353729),
 (694397537353729, 695737600794624),
 (693075572527105, 929811762360322),
 (695737600794624, 700168400654337),
 (929811762360322, 931830115979265),
 (931830115979265, 951912745500672),
 (951912745500672, 965073687117824)]

实现方案

纯Python实现(无需额外依赖)

通过构建父子关系映射+深度优先遍历实现,保证同会话聚合,且根节点、同层级回复均按发布时间排序:

# 1. 构建基础映射
uid_to_item = {uid: (ref, uid) for ref, uid in data}
all_uids = set(uid_to_item.keys())
# 存储每个推文的所有直接回复,按发布时间顺序存入
parent_to_children = {}
for ref, uid in data:
    if ref in all_uids:
        parent_to_children.setdefault(ref, []).append(uid)

processed = set()
sorted_result = []

# 2. 深度优先遍历聚合会话
def dfs(uid):
    if uid in processed:
        return
    # 添加当前推文
    sorted_result.append(uid_to_item[uid])
    processed.add(uid)
    # 递归添加当前推文的所有回复
    for child_uid in parent_to_children.get(uid, []):
        dfs(child_uid)

# 3. 按原始时间顺序遍历,未处理的即为新会话根节点
for ref, uid in data:
    if uid not in processed:
        dfs(uid)

# 输出结果
print(sorted_result)

Pandas实现(可选)

如果需要保留DataFrame格式,可在上述逻辑基础上调整:

# 先拿到排序后的uid顺序
sorted_uids = [item[1] for item in sorted_result]
# 按顺序重排DataFrame
df_sorted = df.set_index('uid').loc[sorted_uids].reset_index()[['reference', 'uid']]

内容的提问来源于stack exchange,提问作者nbego

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 02:54:02