You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取groupby计数占比并绘制各Cluster的双柱对比图?

问题解决方案

1. 将计数转换为占比

提供两种简洁的实现方式:

方式一:利用value_counts的normalize参数

按Cluster分组后,直接计算组内ValueT1的占比:

# 计算每个Cluster下ValueT1为0/1的百分比占比
prob_df = test_df.groupby('Cluster')['ValueT1'].value_counts(normalize=True).mul(100).round(1).reset_index(name='Probability')
  • normalize=True:直接返回0-1区间的占比
  • mul(100):转换为百分比格式
  • round(1):保留1位小数,匹配示例的48.7%、51.2%格式
  • reset_index:将结果转为结构化DataFrame,方便后续绘图

方式二:手动计算占比

若需要更清晰的分步逻辑,可先统计计数再计算占比:

# 统计每个(Cluster, ValueT1)组合的样本数
count_df = test_df.groupby(['Cluster', 'ValueT1'])['Cluster'].count().reset_index(name='Count')
# 统计每个Cluster的总样本数
total_counts = test_df.groupby('Cluster')['Cluster'].count().reset_index(name='Total')
# 合并数据并计算占比
prob_df = count_df.merge(total_counts, on='Cluster')
prob_df['Probability'] = (prob_df['Count'] / prob_df['Total']).mul(100).round(1)

2. 绘制所有Cluster的双柱图

推荐用seaborn实现,代码简洁易读:

依赖安装(未安装时执行)

pip install seaborn matplotlib

绘图代码

import seaborn as sns
import matplotlib.pyplot as plt

# 设置绘图风格
sns.set_style("whitegrid")

# 绘制双柱图
plt.figure(figsize=(10, 6))
sns.barplot(data=prob_df, x='Cluster', y='Probability', hue='ValueT1')

# 添加图表标注
plt.title('每个Cluster下一时刻Value为0/1的概率分布')
plt.xlabel('Cluster编号')
plt.ylabel('概率 (%)')
plt.legend(title='下一时刻Value值')

# 显示图形
plt.show()

若使用pandas原生绘图,可按以下方式实现:

# 将数据转为宽格式(Cluster为行,ValueT1为列)
wide_prob_df = prob_df.pivot(index='Cluster', columns='ValueT1', values='Probability').fillna(0)

# 绘制双柱图
wide_prob_df.plot(kind='bar', figsize=(10, 6))
plt.title('每个Cluster下一时刻Value为0/1的概率分布')
plt.xlabel('Cluster编号')
plt.ylabel('概率 (%)')
plt.legend(title='下一时刻Value值')
plt.show()

内容的提问来源于stack exchange,提问作者Lleims

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 08:50:46