按顺序合并DataFrame中同用户名的连续聊天行
合并DataFrame中同一用户连续发送的聊天记录
我有一个聊天记录DataFrame,需要将同一用户连续发送的多条聊天记录合并为一条,预期效果如下:
| author_username | Content |
|---|---|
| Denise | I want to die so bad. I don’t feel the need to do anything but with an exam coming up, she threw me away like trash. With all the pressure, I don’t want to live. |
| Kenton | Please stay strong, I can feel you. My test just ended next week, back then i feel i don't have hope, and when pandemic first started. I lost contact With all my friends. |
| Denise | Oh |
| Kenton | But look at me now |
| Denise | I cant see you |
| Kenton | ? wdym? |
| Denise | I can't see you |
| Kenton | I know. That is a sentence that people use to make example of themself. So I use that sentence |
| Denise | Ok sry |
之前尝试的几种方法均未达到预期:
- 直接按
author_username分组合并:会把同一用户所有聊天记录合并,无法区分连续发送的批次 - 循环遍历拼接:因pandas的
loc赋值逻辑问题,导致数据未正确更新
解决方案
核心思路是先给连续的同一用户聊天记录标记分组ID,再按分组ID合并内容。
实现代码
import pandas as pd # 假设原始DataFrame为df,包含author_username和content列 # 1. 生成连续用户的分组标识:当前用户与上一行不同时,分组ID自增 df['group_id'] = (df['author_username'] != df['author_username'].shift(1)).cumsum() # 2. 按分组ID合并内容,保留用户名和合并后的聊天内容 merged_df = df.groupby('group_id').agg( author_username=('author_username', 'first'), Content=('content', ' '.join) ).reset_index(drop=True) # 输出结果 print(merged_df)
代码说明
- 生成group_id:通过
shift(1)获取上一行的用户名,与当前行对比,不同则生成新的分组ID,确保连续发送的消息被归为同一组 - 分组合并:按group_id分组,取每组的第一个用户名(同一组内用户一致),将content用空格拼接,最后移除group_id列得到目标结果
内容的提问来源于stack exchange,提问作者Anurag Chaudhary
相关产品推荐
相关产品推荐

