You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Plotly绘制带组内百分比标注的堆叠/分组直方图方法

实现方案

原生px.histogram自带的占比统计是全局维度的,分母为全量样本数,无法按Box分组单独计算组内占比,因此需要先手动完成分箱、组内占比统计,再绘制柱状图(和直方图视觉效果完全一致),同时自定义标注。

  • 第一步:预处理数据,对齐原绘图参数的分箱规则,计算每个Box下各分箱的组内百分比
import pandas as pd
import plotly.express as px
import numpy as np

# 对齐原参数的分箱规则:0-100范围切10个等宽分箱
bin_edges = np.linspace(0, 100, 11)
bin_labels = [f"{int(bin_edges[i])}-{int(bin_edges[i+1])}" for i in range(len(bin_edges)-1)]

# 给每条数据匹配对应分箱
sample_data['value_bin'] = pd.cut(
    sample_data['Value'],
    bins=bin_edges,
    labels=bin_labels,
    include_lowest=True
)

# 统计每个Box的总样本量
box_total = sample_data.groupby('Box')['Value'].count().rename('box_total').reset_index()

# 统计每个Box在各分箱的样本数,计算组内百分比
bin_stat = sample_data.groupby(['Box', 'value_bin'], observed=False)['Value'].count().reset_index(name='count')
bin_stat = bin_stat.merge(box_total, on='Box')
bin_stat['pct'] = bin_stat['count'] / bin_stat['box_total'] * 100
# 匹配直方图的x轴位置
bin_stat['bin_pos'] = bin_stat['value_bin'].apply(lambda x: int(x.split('-')[0]))
  • 第二步:绘制分组/堆叠柱状图,自定义柱体标注为组内百分比
fig = px.bar(
    bin_stat,
    x='bin_pos',
    y='count', # 若需要柱高直接对应百分比,替换为y='pct'
    color='Box',
    barmode='group', # 需要堆叠效果可替换为barmode='stack'
    text=bin_stat['pct'].apply(lambda x: f"{x:.1f}%" if x>0 else ""), # 空分箱不显示标注
    range_x=[0,100]
)

# 调整标注和坐标轴样式,对齐原生直方图效果
fig.update_traces(textposition='outside')
fig.update_xaxes(tickvals=list(range(5,100,10)), ticktext=bin_labels)
fig.show()

效果验证:以提供的示例数据为例

  • Box A总样本量3个,分别落在10-20、20-30、90-100三个分箱,每个分箱标注33.3%
  • Box B总样本量5个,10-20分箱3个样本标注60%,20-30、30-40分箱各1个样本各标注20%
  • Box C总样本量2个,80-90、90-100分箱各1个样本各标注50%

内容的提问来源于stack exchange,提问作者enterML

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.01 12:01:04