You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas高效统计群聊数据中的三方会话数量

高效统计群聊数据中的三方会话数量

需求说明

需要统计群聊数据集中的三方会话数量,规则如下:

  • 三方会话:同一group_id内,必须满足「red用户发消息→green用户回复→red用户再回复」的时间序列
  • Touchpoint定义:
    • Touchpoint1:red用户的第一条消息
    • Touchpoint2:green用户的第一条消息
    • Touchpoint3:red用户在对应Touchpoint2之后发送的第二条消息

示例数据集

import pandas as pd
from pandas import Timestamp

t1_df = pd.DataFrame({'from_red': [True, False, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, False, True], 
              'sent_time': [Timestamp('2021-05-01 06:26:00'), Timestamp('2021-05-04 10:35:00'), Timestamp('2021-05-07 12:16:00'), Timestamp('2021-05-07 12:16:00'), Timestamp('2021-05-09 13:39:00'), Timestamp('2021-05-11 10:02:00'), Timestamp('2021-05-12 13:10:00'), Timestamp('2021-05-12 13:10:00'), Timestamp('2021-05-13 09:46:00'), Timestamp('2021-05-13 22:30:00'), Timestamp('2021-05-14 14:14:00'), Timestamp('2021-05-14 17:08:00'), Timestamp('2021-06-01 09:22:00'), Timestamp('2021-06-01 21:26:00'), Timestamp('2021-06-03 20:19:00'), Timestamp('2021-06-03 20:19:00'), Timestamp('2021-06-09 07:24:00'), Timestamp('2021-05-01 06:44:00'), Timestamp('2021-05-01 08:01:00'), Timestamp('2021-05-01 08:09:00')], 
              'w_uid': ['w_000001', 'w_112681', 'w_002516', 'w_002514', 'w_004073', 'w_005349', 'w_006803', 'w_006804', 'w_008454', 'w_009373', 'w_010063', 'w_010957', 'w_066840', 'w_071471', 'w_081446', 'w_081445', 'w_106472', 'w_000002', 'w_111906', 'w_000003'], 
              'user_id': ['red_00001', 'green_0263', 'red_01071', 'red_01071', 'red_01552', 'red_01552', 'red_02282', 'red_02282', 'red_02600', 'red_02854', 'red_02854', 'red_02600', 'red_00001', 'red_09935', 'red_10592', 'red_10592', 'red_12292', 'red_00002', 'green_0001', 'red_00003'], 
              'group_id': [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1], 
              'touchpoint': [1, 2, 1, 3, 1, 3, 1, 3, 1, 1, 3, 3, 3, 1, 1, 3, 1, 1, 2, 1]}, 
                     columns = ['from_red', 'sent_time', 'w_uid', 'user_id', 'group_id', 'touchpoint'])

t1_df['sent_time'] = pd.to_datetime(t1_df['sent_time'])

数据集结构:

from_redsent_timew_uiduser_idgroup_idtouchpoint
True2021-05-01 06:26:00w_000001red_0000101
False2021-05-04 10:35:00w_112681green_026302
True2021-05-07 12:16:00w_002516red_0107101
True2021-05-07 12:16:00w_002514red_0107103
True2021-05-09 13:39:00w_004073red_0155201
True2021-05-11 10:02:00w_005349red_0155203
True2021-05-12 13:10:00w_006803red_0228201
True2021-05-12 13:10:00w_006804red_0228203
True2021-05-13 09:46:00w_008454red_0260001
True2021-05-13 22:30:00w_009373red_0285401
True2021-05-14 14:14:00w_010063red_0285403
True2021-05-14 17:08:00w_010957red_0260003
True2021-06-01 09:22:00w_066840red_0000103
True2021-06-01 21:26:00w_071471red_0993501
True2021-06-03 20:19:00w_081446red_1059201
True2021-06-03 20:19:00w_081445red_1059203
True2021-06-09 07:24:00w_106472red_1229201
True2021-05-01 06:44:00w_000002red_0000211
False2021-05-01 08:01:00w_111906green_000112
True2021-05-01 08:09:00w_000003red_0000311

现有代码的性能问题

用户提供的代码存在以下低效点:

  1. 循环逻辑错误:遍历sent_time的长度而非唯一group_id,做了大量无效计算
  2. 多次query调用:每次查询都要扫描全量数据,时间复杂度高
  3. 反复pd.concat:每次拼接都会创建新DataFrame,内存开销大且速度慢
  4. 冗余操作:merge(x, how="outer")没有实际作用,属于无效代码

高效实现方案

利用pandas的分组聚合和向量操作,避免循环和重复查询,代码如下:

# 1. 按group_id分组,提取关键信息:是否同时有red/green用户、最早的touchpoint2时间
group_stats = t1_df.groupby('group_id').agg(
    has_both_users=('from_red', lambda x: x.nunique() == 2),
    first_touchpoint2_time=('sent_time', lambda x: x[t1_df.loc[x.index, 'touchpoint'] == 2].min())
).reset_index()

# 2. 筛选出符合条件的群组:同时有red/green,且存在touchpoint2
valid_groups = group_stats[group_stats['has_both_users'] & group_stats['first_touchpoint2_time'].notna()]

# 3. 关联原数据,筛选出符合时间要求的touchpoint3记录
valid_three_party = t1_df.merge(valid_groups, on='group_id').query(
    "touchpoint == 3 and sent_time > first_touchpoint2_time"
)

# 4. 统计三方会话数量
three_party_total = len(valid_three_party)
print(f"三方会话总数:{three_party_total}")

# 查看具体会话记录
print(valid_three_party)

方案优势

  • 向量操作替代循环:全程使用pandas的向量化方法,效率比循环提升数倍甚至数十倍
  • 单次聚合获取群组信息:一次分组聚合就拿到所有群组的关键指标,避免多次查询
  • 避免冗余拼接:直接通过merge和筛选得到结果,没有内存浪费的concat操作
  • 逻辑更准确:针对group_id处理,而非错误遍历sent_time长度

内容的提问来源于stack exchange,提问作者n3a5p7s9t1e3r

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 06:15:40