pandas按组绘制变量均值柱状图时如何标注每组实例数量
实现方法
核心是先把分组的均值、样本数两个指标同时聚合出来,再根据展示需求选对应方案即可。
第一步:提前聚合分组统计结果
不要直接链式调用plot,先把分组计算结果存为变量,方便后续调用:
# 按Borrow_Rank分组,同时计算Outcome均值、每组样本行数 group_stats = df.groupby('Borrow_Rank', as_index=False).agg( outcome_mean = ('Outcome', 'mean'), sample_cnt = ('Outcome', 'count') )
方案1:柱子顶部标注样本数(推荐)
柱子高度保持0-1的Outcome均值,直接在每个柱子上方加文字标注对应样本量,不会因为量纲差异造成视觉误导,可读性最好:
import matplotlib.pyplot as plt # 绘制基础均值柱状图 ax = group_stats.plot( x='Borrow_Rank', y='outcome_mean', kind='bar', figsize=(9, 5), color='#2c7fb8', ylim=(0, 1.1), # 顶部预留10%空间放标注,避免文字被截断 xlabel='Borrow Rank', ylabel='Mean Outcome' ) # 循环给每个柱子加样本数标注 for bar_idx, row in group_stats.iterrows(): ax.text( x=bar_idx, y=row['outcome_mean'] + 0.02, # 标注位置在柱子顶部上方2%高度处 s=f"n={int(row['sample_cnt'])}", ha='center', fontsize=9 ) plt.xticks(rotation=0) plt.tight_layout() plt.show()
方案2:双Y轴展示两组柱形
如果你需要直观对比不同分组的样本量差异,可以用双Y轴分别承载均值(0-1)、样本数(2000-5000)两个量纲差异大的指标,注意错开柱子位置、用不同颜色区分:
import matplotlib.pyplot as plt fig, ax_mean = plt.subplots(figsize=(9,5)) # 左Y轴画Outcome均值 group_stats.plot( x='Borrow_Rank', y='outcome_mean', kind='bar', width=0.35, position=1, color='#2c7fb8', ax=ax_mean, xlabel='Borrow Rank', ylabel='Mean Outcome', ylim=(0, 1.1) ) # 右Y轴画样本计数 ax_cnt = ax_mean.twinx() group_stats.plot( x='Borrow_Rank', y='sample_cnt', kind='bar', width=0.35, position=0, color='#ff7f0e', ax=ax_cnt, ylabel='Sample Count', ylim=(1800, 5200) # 比实际2000-5000的区间稍宽,预留边距 ) # 合并两个轴的图例,避免重复 h1, l1 = ax_mean.get_legend_handles_labels() h2, l2 = ax_cnt.get_legend_handles_labels() ax_mean.legend(h1+h2, ['Mean Outcome', 'Sample Count'], loc='upper right') ax_cnt.get_legend().remove() plt.xticks(rotation=0) plt.tight_layout() plt.show()
注意:双轴方案容易让读者误读柱形高度的对应关系,除非明确需要展示样本数的组间差异,否则优先选择方案1的标注形式。
内容的提问来源于stack exchange,提问作者simply_here
相关产品推荐
相关产品推荐

