聊天分析中如何将Pandas DataFrame特定行值合并求和为一行?
解决方法
方法1:指定用户合并为Other
直接定义需要合并的用户列表,通过映射分组后求和:
import pandas as pd # 原始数据构造(替换成你的实际DataFrame) data = { '用户名': ['A', 'B', 'C', 'D', 'E', 'F', 'G', 'H', 'I', 'J'], '消息数量': [234, 128, 112, 97, 86, 43, 32, 24, 22, 9] } df = pd.DataFrame(data) # 定义需要合并到Other的用户 other_users = ['F', 'G', 'H', 'I', 'J'] # 创建分组映射:目标用户标记为Other,其余保留原名 df['分组'] = df['用户名'].map(lambda x: 'Other' if x in other_users else x) # 按分组求和并整理结构 result_df = df.groupby('分组')['消息数量'].sum().reset_index().rename(columns={'分组': '用户名'}) print(result_df)
运行后输出结果:
用户名 消息数量 0 A 234 1 B 128 2 C 112 3 D 97 4 E 86 5 Other 130
方法2:自动筛选合并(更灵活,适合饼图场景)
如果需要动态合并消息量少的用户(比如保留前N个活跃用户,剩余合并),可以用以下方式:
# 保留消息量前5的用户,其余合并为Other top_n = 5 # 获取前N个用户名列表 top_users = df.nlargest(top_n, '消息数量')['用户名'].tolist() # 映射分组 df['分组'] = df['用户名'].map(lambda x: x if x in top_users else 'Other') result_df = df.groupby('分组')['消息数量'].sum().reset_index().rename(columns={'分组': '用户名'})
也可以按消息量阈值筛选(比如合并消息数低于90的用户):
threshold = 90 df['分组'] = df.apply(lambda row: 'Other' if row['消息数量'] < threshold else row['用户名'], axis=1) result_df = df.groupby('分组')['消息数量'].sum().reset_index().rename(columns={'分组': '用户名'})
内容的提问来源于stack exchange,提问作者Vayun Ekbote
相关产品推荐
相关产品推荐

