如何使用pandas简便计算按地区分组的加权平均分数?
方法一:全向量化运算(推荐用于大型数据集)
直接通过逐行乘积+分组求和计算,全程无循环,性能最优:
# 计算各得分项与对应权重的乘积 df[["score1_weighted", "score2_weighted"]] = df[["score1", "score2"]].mul(df["weight"], axis=0) # 按地区分组聚合 grouped = df.groupby("region", as_index=False).agg( sum_score1=("score1_weighted", "sum"), sum_score2=("score2_weighted", "sum"), sum_weight=("weight", "sum") ) # 计算加权平均得到最终结果 result = grouped.assign( score1_wavg = grouped["sum_score1"] / grouped["sum_weight"], score2_wavg = grouped["sum_score2"] / grouped["sum_weight"] )[["region", "score1_wavg", "score2_wavg"]]
运行后得到的result就是你需要的2行3列格式的输出。
方法二:内置函数快速实现(代码更简洁)
直接用numpy自带的加权平均函数np.average配合分组apply,代码量更小:
result = df.groupby("region", as_index=False).apply( lambda group: pd.Series({ "score1_wavg": np.average(group["score1"], weights=group["weight"]), "score2_wavg": np.average(group["score2"], weights=group["weight"]) }) )
如果你的数据集规模很大,优先选择第一种方法,向量化运算的性能远高于逐组迭代的apply逻辑。
内容的提问来源于stack exchange,提问作者Cooper
相关产品推荐
相关产品推荐

