You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取非计数轴条形图的分组样本数量?

获取Seaborn分组条形图的样本计数并标注

我用Seaborn绘制了分组条形图,x轴为day,y轴为tip,通过hue="sex"按性别分组。现在需要获取每个条形对应的样本计数(比如周日付小费的男性有8人、女性有1人),并标注在图表上。目前我只会标注条形的高度值,但不知道如何获取这些样本计数值。

原代码:

import seaborn as sns
import matplotlib.pyplot as plt

sns.set_theme(style="whitegrid")

df = sns.load_dataset("tips")[1:10] 
print(df)

ax = sns.barplot(x='day', y='tip',hue="sex", data=df, palette="tab20_r")

for rect in ax.patches:
    y_value = rect.get_height()
    x_value = rect.get_x() + rect.get_width() / 2
    space = 1
    label = "{:.0f}".format(y_value)
    ax.annotate(label, (x_value, y_value), xytext=(0, space), textcoords="offset points", ha='center', va='bottom')
plt.show()

解决方案

要获取每个分组的样本计数,先对数据按day和sex分组统计,再将统计结果与条形图的每个条形对应标注:

  1. 统计分组样本数:用groupby计算每个(日期,性别)组合的样本量
  2. 匹配条形顺序:按条形图的x轴、hue类别顺序整理计数,确保和绘制顺序一致
  3. 标注计数:将计数逐个对应到条形上

修改后的完整代码:

import seaborn as sns
import matplotlib.pyplot as plt

sns.set_theme(style="whitegrid")

df = sns.load_dataset("tips")[1:10] 
print(df)

# 统计每个(day, sex)组的样本数量
count_data = df.groupby(['day', 'sex']).size().reset_index(name='count')

# 绘制条形图
ax = sns.barplot(x='day', y='tip', hue="sex", data=df, palette="tab20_r")

# 获取条形图的x轴类别顺序和hue类别顺序
x_categories = [tick.get_text() for tick in ax.get_xticklabels()]
hue_categories = ax.get_legend_handles_labels()[1]

# 按条形绘制顺序整理计数列表
count_list = []
for day in x_categories:
    for sex in hue_categories:
        # 查找对应组的计数,无数据则记为0
        cnt = count_data[(count_data['day'] == day) & (count_data['sex'] == sex)]['count'].values
        count_list.append(cnt[0] if len(cnt) > 0 else 0)

# 给每个条形标注样本计数
for rect, cnt in zip(ax.patches, count_list):
    x_pos = rect.get_x() + rect.get_width() / 2
    y_pos = rect.get_height()
    # 标注在条形顶部,可调整space控制距离
    ax.annotate(f'N={cnt}',
                (x_pos, y_pos),
                xytext=(0, 0.3),
                textcoords='offset points',
                ha='center', va='bottom',
                fontsize=9)

plt.show()

说明

  • groupby(['day', 'sex']).size()会精准统计每个分组的样本数量,reset_index把结果转为DataFrame方便查找。
  • 通过ax.get_xticklabels()和ax.get_legend_handles_labels()获取条形的绘制顺序,确保计数和条形一一对应,避免错位。
  • 如果需要同时显示tip均值和计数,可以调整标注位置(比如把计数放在条形内部)。

内容的提问来源于stack exchange,提问作者Mohit Narwani

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 10:45:47