You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas按payment和country分组统计新增唯一邮箱均值指标如何实现

解决方案

你可以在原有代码的基础上,额外统计每个分组下的唯一邮箱数量,再将已计算好的各类求和指标除以该数量,即可得到你需要的人均指标:

import pandas as pd

# 原始DataFrame
df = pd.DataFrame({'country':['AU','DE','AU','US','FR','FR'],
                   'payment':['visa','paypal','visa','visa','visa','visa'], 
                   'type':['type1','type1','type2','type2','type1','type1'],
                   'err':['OK','OK','ERROR','OK','OK','ERROR'],
                   'email': ['user1@email.com','user2@email.com','user1@email.com','user4@email.com','user2@email.com','user2@email.com'] })

# 原有逻辑:分组前的数据集新增email列,方便后续统计唯一邮箱数
c = df['err'].eq("ERROR")
g = (df[['payment','country', 'email']]
     .assign(num_errors=c,**pd.get_dummies(df[['type']],prefix=['num']))
     .groupby(['payment','country']))
out = g.size().to_frame("number_payments").join(g.sum(numeric_only=True)).reset_index()

# 补全原数据中不存在的num_type3
out['num_type3'] = 0

# 统计每个(payment, country)分组下的唯一邮箱数量
unique_email_cnt = df.groupby(['payment','country'])['email'].nunique()
out = out.merge(unique_email_cnt.rename('unique_email_cnt'), on=['payment','country'])

# 计算各类按唯一邮箱平均的指标
out['num_errors_per_unique_email'] = out['num_errors'] / out['unique_email_cnt']
out['num_type1_per_unique_email'] = out['num_type1'] / out['unique_email_cnt']
out['num_type2_per_unique_email'] = out['num_type2'] / out['unique_email_cnt']
out['num_type3_per_unique_email'] = out['num_type3'] / out['unique_email_cnt']

# 删除不需要的中间列,得到最终结果
out = out.drop(columns='unique_email_cnt')

如果需要平均列输出为整数,可在计算后添加.astype(int)做类型转换即可。

内容的提问来源于stack exchange,提问作者Alex_Y

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 00:51:02