Plotly Python go.Histogram添加分箱占比悬停与柱上标注方法
Plotly 叠加直方图添加分箱占比悬停与柱顶标注实现
修改后完整代码
import plotly.graph_objects as go import numpy as np lapsed_df = claim_df[claim_df['EarlyClaimAndLapsed']=='Y'] non_lapsed_df = claim_df[claim_df['EarlyClaimAndLapsed']=='N'] # 固定分箱数量,和原代码逻辑保持一致 BIN_COUNT = 10 for con_col in claim_df.select_dtypes(include = ['number']).columns.tolist(): # 1. 基于全列数据计算统一分箱边界,保证两组分箱完全对齐 col_all = claim_df[con_col].dropna() _, bin_edges = np.histogram(col_all, bins=BIN_COUNT) bin_size = bin_edges[1] - bin_edges[0] bin_mids = 0.5 * (bin_edges[:-1] + bin_edges[1:]) # 计算分箱中点,用于放置柱顶标注 # 2. 分别计算两组在统一分箱规则下的计数、占比 n_counts = np.histogram(non_lapsed_df[con_col].dropna(), bins=bin_edges)[0] y_counts = np.histogram(lapsed_df[con_col].dropna(), bins=bin_edges)[0] total_counts = n_counts + y_counts # 处理分箱无数据时的除零异常 n_pct = np.where(total_counts == 0, 0, n_counts / total_counts * 100) y_pct = np.where(total_counts == 0, 0, y_counts / total_counts * 100) fig = go.Figure() # 添加未失效组(N)直方图 fig.add_trace(go.Histogram( x = non_lapsed_df[con_col], name='N', marker_color='blue', customdata=n_pct, # 传入预计算的占比数据供悬停调用 xbins=dict( start=bin_edges[0], end=bin_edges[-1], size=bin_size ) )) # 添加失效组(Y)直方图 fig.add_trace(go.Histogram( x = lapsed_df[con_col], name='Y', marker_color='red', customdata=y_pct, # 传入预计算的占比数据供悬停调用 xbins=dict( start=bin_edges[0], end=bin_edges[-1], size=bin_size ) )) # 3. 配置基础布局 fig.update_layout( title=f'{con_col} vs EarlyClaimAndLapsed', title_x=0.5, barmode="overlay", legend_title='EarlyClaimAndLapsed', xaxis_title=f'{con_col}', yaxis_title='Count' ) # 4. 配置悬停模板,新增分箱占比字段 fig.update_traces( opacity=0.5, histfunc='count', hovertemplate="<br>".join([ f"{con_col}"+": %{x}", "Count: %{y}", "EarlyClaimAndLapsed: %{name}", "分箱占比: %{customdata:.2f}%", "<extra></extra>" # 移除悬停框默认冗余侧边信息 ]) ) # 5. 添加柱顶占比标注 for i in range(len(bin_mids)): # 标注N组占比 if n_counts[i] > 0: fig.add_annotation( x=bin_mids[i], y=n_counts[i], text=f"{n_pct[i]:.2f}%", showarrow=False, yshift=5, # 向上偏移避免和柱体重合 font=dict(color='blue', size=10) ) # 标注Y组占比 if y_counts[i] > 0: fig.add_annotation( x=bin_mids[i], y=y_counts[i], text=f"{y_pct[i]:.2f}%", showarrow=False, yshift=5, font=dict(color='red', size=10) ) fig.show()
关键修改说明
- 统一分箱规则:用numpy对字段全量数据生成分箱边界,强制两组直方图使用完全一致的分箱区间,避免分箱错位导致占比计算错误。
- 预计算占比数据:基于两组的分箱计数,按照给定公式计算每个分箱下两组的占比,通过
customdata参数传入直方图trace供悬停模板调用,同时处理了空分箱的除零异常。 - 自定义悬停逻辑:在原有悬停字段后新增分箱占比项,保留两位小数显示为百分比格式,同时移除默认多余的悬停侧边信息。
- 柱顶标注实现:遍历所有非空分箱,在对应柱顶位置添加和柱体同色的占比文本,设置小幅向上偏移避免遮挡柱体。
- 可按需调整
yshift偏移量、标注字体大小、分箱数量等参数适配实际显示效果。
内容的提问来源于stack exchange,提问作者Ming Jun Lim
相关产品推荐
相关产品推荐

