如何使用seaborn为分类数据绘制100%堆叠条形图
实现方案
前置数据预处理
先计算每个Rank分组下不同Clicked分类的占比,得到适配绘图的宽格式占比表:
import pandas as pd import matplotlib.pyplot as plt import seaborn as sns # 计算各Rank下的总样本量 rank_total = df.groupby('Rank')['Clicked'].count() # 计算各Rank下不同Clicked分类的样本量 cat_count = df.groupby(['Rank', 'Clicked'])['Clicked'].count().unstack(fill_value=0) # 计算占比 ratio_df = cat_count.div(rank_total, axis=0) # 固定分类顺序为Cat1、Cat2、Cat3、Cat4,避免乱序 ratio_df = ratio_df.reindex(['Cat1', 'Cat2', 'Cat3', 'Cat4'], axis=1, fill_value=0)
方案1:纯Matplotlib实现
通过Matplotlib原生的bar接口逐层堆叠绘制:
x = range(len(ratio_df.index)) # 堆叠的底部起始位置,初始为0 bottom = [0]*len(ratio_df.index) # 自定义配色,可按需替换 colors = sns.color_palette('tab10', n_colors=4) for idx, cat in enumerate(ratio_df.columns): plt.bar( x=x, height=ratio_df[cat], bottom=bottom, label=cat, color=colors[idx] ) # 更新下一层的堆叠底部位置 bottom = [bottom[i] + ratio_df.iloc[i, idx] for i in range(len(bottom))] plt.xticks(x, ratio_df.index) plt.xlabel('Rank') plt.ylabel('占比') plt.legend() plt.show()
方案2:Seaborn辅助实现
Seaborn无原生堆叠条形图接口,可将数据转为长格式后用histplot的fill模式实现:
# 宽表转长格式适配Seaborn接口 long_df = ratio_df.reset_index().melt( id_vars='Rank', var_name='Clicked', value_name='ratio' ) sns.histplot( data=long_df, x='Rank', weights='ratio', hue='Clicked', multiple='fill', shrink=0.8 # 调整柱子宽度,匹配pandas plot的默认效果 ) plt.xlabel('Rank') plt.ylabel('占比') plt.show()
内容的提问来源于stack exchange,提问作者amestrian
相关产品推荐
相关产品推荐

