基于多级索引Pandas DataFrame绘制按场景分色的Seaborn箱线图
解决方案
1. 重置多级索引解决报错
Could not interpret input 'season'报错的直接原因是season和period是DataFrame的索引,不在普通列中,首先执行索引重置操作:
# 将period、season两级索引转为可直接调用的普通列 df = df.reset_index()
2. 适配预计算统计量的绘图实现
默认sns.boxplot是基于原始样本自动计算中位数、四分位距等统计量,你已经提前完成了均值、标准差、最值的计算,以下给出两种可行绘图方案:
方案1:Matplotlib bxp 自定义统计量箱线图
该方案可以完全匹配你已有的预计算结果,精准定义箱体各个统计节点:
import matplotlib.pyplot as plt import seaborn as sns import pandas as pd # 重置索引 df = df.reset_index() # 生成x轴分类标签:period + season 组合 df['x_category'] = df['period'].astype(str) + '_' + df['season'].astype(str) # 按场景分组配置调色盘 scenarios = df['model_scenario'].unique() palette = sns.color_palette("Set2", n_colors=len(scenarios)) scenario_color_map = dict(zip(scenarios, palette)) fig, ax = plt.subplots(figsize=(12, 6)) bxp_stats = [] positions = [] box_colors = [] x_cats = sorted(df['x_category'].unique()) # 同分类下不同场景箱体偏移配置,避免重叠 offset_step = 0.2 base_pos = list(range(len(x_cats))) for cat_idx, cat in enumerate(x_cats): cat_subset = df[df['x_category'] == cat] for scen_idx, scen in enumerate(scenarios): scen_row = cat_subset[cat_subset['model_scenario'] == scen].iloc[0] # 构造bxp要求的统计量格式 stats = { 'med': scen_row['hedges'], 'q1': scen_row['hedges'] - scen_row['hedges_std'], 'q3': scen_row['hedges'] + scen_row['hedges_std'], 'whislo': scen_row['hedges_min'], 'whishi': scen_row['hedges_max'], 'fliers': [] } bxp_stats.append(stats) positions.append(cat_idx - offset_step*(len(scenarios)/2) + scen_idx*offset_step) box_colors.append(scenario_color_map[scen]) # 绘制箱线图 bxp_result = ax.bxp(bxp_stats, positions=positions, patch_artist=True, showfliers=False) # 配置箱体颜色 for patch, color in zip(bxp_result['boxes'], box_colors): patch.set_facecolor(color) # 坐标轴与图例配置 ax.set_xticks(base_pos) ax.set_xticklabels(x_cats, rotation=45) ax.set_xlabel('Period_Season') ax.set_ylabel('Hedges Estimate') from matplotlib.patches import Patch legend_elements = [Patch(facecolor=scenario_color_map[s], label=s) for s in scenarios] ax.legend(handles=legend_elements, title='Model Scenario') plt.tight_layout() plt.show()
方案2:生成模拟样本适配seaborn boxplot
如果不需要严格绑定预计算的统计值,可通过生成模拟样本快速实现seaborn风格绘图:
import numpy as np import seaborn as sns import matplotlib.pyplot as plt df = df.reset_index() df['x_category'] = df['period'].astype(str) + '_' + df['season'].astype(str) # 基于统计量生成模拟样本 sim_records = [] for _, row in df.iterrows(): samples = np.random.normal(loc=row['hedges'], scale=row['hedges_std'], size=100) samples = np.clip(samples, row['hedges_min'], row['hedges_max']) for val in samples: sim_records.append({ 'x_category': row['x_category'], 'model_scenario': row['model_scenario'], 'hedges': val }) sim_df = pd.DataFrame(sim_records) # 调用seaborn绘制箱线图 plt.figure(figsize=(12, 6)) sns.boxplot(data=sim_df, x='x_category', y='hedges', hue='model_scenario', palette='Set2') plt.xticks(rotation=45) plt.xlabel('Period_Season') plt.ylabel('Hedges Estimate') plt.tight_layout() plt.show()
内容的提问来源于stack exchange,提问作者Trond Kristiansen
相关产品推荐
相关产品推荐

