如何在Pandas中分组统计并保留Churn列及计数
解决方案
要实现你需要的统计效果,核心是按Account、Booking、Churn三个字段联合分组,再对每组的记录数进行统计,最后处理Account列的重复显示即可。
步骤1:构造测试数据(模拟你的数据集)
import pandas as pd data = { 'Account': ['ABC inc', 'ABC inc', 'ABC inc', 'Company A', 'Company A', 'Company A'], 'Booking': ['New', 'Upsell', 'New', 'Renew', 'New', 'Renew'], 'Churn': [1, 0, 1, 0, 1, 1] } df = pd.DataFrame(data)
步骤2:分组统计记录数
通过groupby同时指定三个分组字段,用size()统计每组的记录数量,再通过reset_index()将分组索引转为普通列,并将计数列命名为Count:
# 分组统计 result = df.groupby(['Account', 'Booking', 'Churn'], as_index=False).size().rename(columns={'size': 'Count'})
步骤3:处理Account列的重复显示
将同一Account的后续行设置为空字符串,让结果和你期望的格式一致:
# 重复的Account值设为空 result['Account'] = result['Account'].where(result['Account'] != result['Account'].shift(), '')
最终结果
执行上述代码后,result的输出如下:
| Account | Booking | Churn | Count |
|---|---|---|---|
| ABC inc | New | 1 | 2 |
| Upsell | 0 | 1 | |
| Company A | Renew | 0 | 1 |
| New | 1 | 1 | |
| Renew | 1 | 1 |
你之前代码的问题说明
你之前仅按ACCOUNT_NAME单字段分组,且使用agg(['unique','nunique']),这只能统计每个Account下Booking的唯一值和唯一值数量,无法关联Churn字段的分组统计,自然得不到你需要的结果。
内容的提问来源于stack exchange,提问作者Austin Meredith
相关产品推荐
相关产品推荐

