You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何计算DataFrame中各品牌按年度的数量占比?

问题描述

假设我的DataFrame如下:

Build yearBrand
2010Mercedes
2010Mercedes
2010BMW
2010Kia
2011Toyota
2011Mercedes
2011Mercedes
2012Tesla

我需要找出Build year与Brand的所有唯一组合,统计每组的数量,同时计算各品牌在对应年份的占比。目前我写了这段代码:

df.groupby(["Build year", "Brand"]).count()

有没有简便方法把结果转换成年度占比?期望输出如下:

Build yearBrandCountPercentage of annual count
2010Mercedes20.5
2010BMW10.25
2010Kia10.25
2011Toyota10.33
2011Mercedes20.66
2012Tesla11
解决方案

可以通过分组统计总数结合groupby.transform快速实现,步骤如下:

  1. 先按Build year和Brand分组统计行数,生成Count列:
result = df.groupby(["Build year", "Brand"]).size().reset_index(name="Count")

用size()比count()更直接,因为我们仅需统计每组的行数。

  1. 计算每个年份的总数量,再用每组的Count除以对应年份的总数得到占比:
result["Percentage of annual count"] = result["Count"] / result.groupby("Build year")["Count"].transform("sum")

transform("sum")会为每行返回其所属年份的总数量,直接做除法就能得到占比。

  1. 可选:如果需要保留两位小数,用round()处理:
result["Percentage of annual count"] = result["Percentage of annual count"].round(2)

完整代码整合如下:

# 统计分组数量
result = df.groupby(["Build year", "Brand"]).size().reset_index(name="Count")
# 计算年度占比并保留两位小数
result["Percentage of annual count"] = (result["Count"] / result.groupby("Build year")["Count"].transform("sum")).round(2)

执行后即可得到符合预期的输出结果。

内容的提问来源于stack exchange,提问作者Xtiaan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 16:50:26