如何基于三特征拆分Seaborn Countplot可视化图表?
问题:在时间框架频率图中按Condition拆分柱子并叠加Outcome占比
需求说明:
需要生成不同时间框架下的频率图,每个时间框架包含4种Condition(A+、A-、B+、B-),且每个Condition的柱子需按High和Low两种Outcome的占比拆分填充(比如4hrs的A+共3个样本,1个High、2个Low,柱子按1/3和2/3比例用深浅色填充)。当前代码仅按Condition拆分柱子,未体现Outcome的比例分布。
原始代码:
import pandas as pd import seaborn as sns import matplotlib.pyplot as plt exp = {'Time':['2hrs', '2hrs', '2hrs','2hrs','2hrs', '2hrs', '2hrs','4hrs','4hrs','4hrs','4hrs','4hrs','4hrs','4hrs','4hrs' ], 'Condition':['A+','A-','B+','B-','B+','B-','B+','A+','A-','B+','B-','A+','A-','A+','A-',], 'Outcome': ['High', 'Low','High','High', 'Low','High', 'Low','High','High','High','Low', 'Low','Low','Low', 'High']} df = pd.DataFrame(data=exp) df.head() hue_order = ['A+', 'A-', 'B+', 'B-'] ax = sns.countplot(data=df, x='Time' , hue='Condition', hue_order=hue_order, palette='Set1') plt.legend(title='', loc='upper left', bbox_to_anchor=(1,1)) plt.show()
解决方案:预处理数据+堆叠柱状图实现占比拆分
要实现每个Condition柱子内按Outcome占比拆分,需要先对数据进行分组统计,再用堆叠柱状图呈现。步骤如下:
1. 数据预处理:统计各分组样本数量
先按Time、Condition、Outcome分组计数,再整理成便于绘图的格式:
import pandas as pd import seaborn as sns import matplotlib.pyplot as plt # 原始数据 exp = {'Time':['2hrs', '2hrs', '2hrs','2hrs','2hrs', '2hrs', '2hrs','4hrs','4hrs','4hrs','4hrs','4hrs','4hrs','4hrs','4hrs' ], 'Condition':['A+','A-','B+','B-','B+','B-','B+','A+','A-','B+','B-','A+','A-','A+','A-',], 'Outcome': ['High', 'Low','High','High', 'Low','High', 'Low','High','High','High','Low', 'Low','Low','Low', 'High']} df = pd.DataFrame(data=exp) # 分组统计:Time-Condition-Outcome的样本数,缺失组合填充0 count_df = df.groupby(['Time', 'Condition', 'Outcome']).size().unstack(fill_value=0) count_df = count_df.reset_index()
2. 绘制堆叠柱状图
设置Seaborn风格,遍历每个Condition在对应Time位置绘制堆叠的High/Low柱子:
# 设置Seaborn风格 sns.set_style("whitegrid") # 定义核心参数 time_order = ['2hrs', '4hrs'] condition_order = ['A+', 'A-', 'B+', 'B-'] # 为每个Condition分配深浅同色系:[High亮色, Low深色] color_palette = { 'A+': ['#e41a1c', '#b30000'], 'A-': ['#377eb8', '#005086'], 'B+': ['#4daf4a', '#2d862d'], 'B-': ['#984ea3', '#7a2e93'] } bar_width = 0.2 # 单个Condition柱子的宽度 # 创建画布 fig, ax = plt.subplots(figsize=(10, 6)) # 遍历每个时间框架,绘制对应Condition的堆叠柱子 for i, time in enumerate(time_order): # 筛选当前时间框架的数据并按Condition顺序排序 time_data = count_df[count_df['Time'] == time] time_data = time_data.set_index('Condition').reindex(condition_order).reset_index() bottom = 0 # 堆叠柱子的底部起始位置 for j, cond in enumerate(condition_order): # 获取当前Condition的High/Low样本数 high_count = time_data.loc[time_data['Condition'] == cond, 'High'].values[0] low_count = time_data.loc[time_data['Condition'] == cond, 'Low'].values[0] # 计算柱子的x轴位置:时间框架基础位置 + Condition偏移量 x_pos = i + j * bar_width # 绘制Low柱子(底层) ax.bar(x_pos, low_count, width=bar_width, color=color_palette[cond][1]) # 绘制High柱子(堆叠在Low之上) ax.bar(x_pos, high_count, width=bar_width, bottom=low_count, color=color_palette[cond][0]) # 可选:在柱子顶部标注总样本数 total = high_count + low_count ax.text(x_pos, total + 0.1, str(total), ha='center', fontsize=10) # 设置坐标轴与标题 ax.set_xticks([i + 1.5*bar_width for i in range(len(time_order))]) ax.set_xticklabels(time_order) ax.set_xlabel('Time Frame') ax.set_ylabel('Sample Count') ax.set_title('Sample Distribution by Time, Condition and Outcome') # 创建自定义图例 handles = [] labels = [] for cond in condition_order: handles.append(plt.Rectangle((0,0),1,1, color=color_palette[cond][0])) labels.append(f'{cond} - High') handles.append(plt.Rectangle((0,0),1,1, color=color_palette[cond][1])) labels.append(f'{cond} - Low') ax.legend(handles, labels, title='Condition & Outcome', loc='upper left', bbox_to_anchor=(1,1)) # 调整布局防止图例截断 plt.tight_layout() plt.show()
代码说明
- 数据预处理:通过
groupby和unstack快速统计各分组样本数,确保缺失的组合自动填充0,避免绘图时出现断层。 - 堆叠绘制:利用
matplotlib的bar函数,通过bottom参数控制堆叠层级,同一个Condition的High/Low使用深浅同色系,保证视觉上的分组一致性。 - 刻度与图例:手动计算x轴刻度位置,让每个时间框架下的4个Condition柱子排列整齐;自定义图例清晰展示每个颜色对应的分组关系。
内容的提问来源于stack exchange,提问作者Lucho
相关产品推荐
相关产品推荐

