求助:使用Matplotlib基于多列分类绘制分组柱状图
基于分类列绘制分组柱状图的解决方案
问题分析
原代码的核心问题:
- X轴坐标计算混乱,导致柱子重叠或位置错位
- 循环中重复添加相同图例,造成图例冗余
- 未合理实现
Type的展示逻辑,也没有适配大规模数据集的通用处理方式
方案1:将Type拆分为子图绘制(更清晰,适合多分类场景)
把不同Type放在独立子图,每个子图展示对应Ratio和Method的分组柱形,代码无需硬编码类别值,适配任意规模数据集。
import pandas as pd import matplotlib.pyplot as plt import numpy as np # 示例数据 df = pd.DataFrame({'Type': ['A','A','A','A','A','A','A','A','A','B','B','B','B','B','B','B','B','B'], 'Ratio': [3, 3, 3, 5, 5, 5, 7, 7, 7,3, 3, 3, 5, 5, 5, 7, 7, 7], 'Method': ['X','Y','Z','X','Y','Z','X','Y','Z','X','Y','Z','X','Y','Z','X','Y','Z'], 'Result': [90, 85, 96, 89, 82, 80, 78, 72, 75, 91, 82, 94, 87, 86, 84, 71, 78, 86]}) # 自动获取所有唯一分类值,无需硬编码 types = df['Type'].unique() ratios = df['Ratio'].unique() methods = df['Method'].unique() n_methods = len(methods) bar_width = 0.25 # 每组内单个柱子的宽度 # 创建子图,数量与Type的类别数一致 fig, axes = plt.subplots(nrows=1, ncols=len(types), figsize=(10, 4), sharey=True) # 循环处理每个Type for ax, t in zip(axes, types): df_type = df[df['Type'] == t] # 生成每个Ratio组的X轴基准位置 x = np.arange(len(ratios)) # 循环绘制每个Method对应的柱子 for i, m in enumerate(methods): # 提取当前Type+Method的Result数据,按Ratio分组取均值(适配多重复值场景) df_sub = df_type[df_type['Method'] == m].groupby('Ratio')['Result'].mean().reindex(ratios) # 计算柱子的X位置:基准位置 + 偏移量,避免重叠 ax.bar(x + i*bar_width, df_sub.values, width=bar_width, label=m) # 设置子图属性 ax.set_title(f'Type: {t}') ax.set_xlabel('Ratio') ax.set_xticks(x + bar_width*(n_methods-1)/2) ax.set_xticklabels(ratios) ax.legend(title='Method') # 统一设置Y轴标签 axes[0].set_ylabel('Result') plt.tight_layout() plt.savefig('grouped_bar_subplots.svg')
方案2:将Type和Ratio合并为X轴分组(同图展示)
如果要把所有内容放在一张图,可将Type与Ratio组合为X轴的分组单元:
import pandas as pd import matplotlib.pyplot as plt import numpy as np df = pd.DataFrame({'Type': ['A','A','A','A','A','A','A','A','A','B','B','B','B','B','B','B','B','B'], 'Ratio': [3, 3, 3, 5, 5, 5, 7, 7, 7,3, 3, 3, 5, 5, 5, 7, 7, 7], 'Method': ['X','Y','Z','X','Y','Z','X','Y','Z','X','Y','Z','X','Y','Z','X','Y','Z'], 'Result': [90, 85, 96, 89, 82, 80, 78, 72, 75, 91, 82, 94, 87, 86, 84, 71, 78, 86]}) # 创建组合分组键:Type + Ratio df['Group'] = df['Type'] + '_' + df['Ratio'].astype(str) groups = df['Group'].unique() methods = df['Method'].unique() n_methods = len(methods) bar_width = 0.2 x = np.arange(len(groups)) fig, ax = plt.subplots(figsize=(12, 4)) # 循环绘制每个Method的柱子 for i, m in enumerate(methods): df_sub = df[df['Method'] == m].groupby('Group')['Result'].mean().reindex(groups) ax.bar(x + i*bar_width, df_sub.values, width=bar_width, label=m) # 设置图表属性 ax.set_xlabel('Type-Ratio Group') ax.set_ylabel('Result') ax.set_xticks(x + bar_width*(n_methods-1)/2) ax.set_xticklabels(groups, rotation=45) ax.legend(title='Method') plt.tight_layout() plt.savefig('grouped_bar_single.svg')
关键优化点
- 避免硬编码:通过
unique()自动获取分类值,适配任意规模的数据集 - X轴位置计算:用
numpy生成基准位置,按Method数量计算偏移,确保柱子排列整齐无重叠 - 数据聚合:使用
groupby+mean处理重复数据(若数据中每个Type-Ratio-Method组合有多条记录,取均值更合理) - 图例控制:每个子图或单图仅添加一次图例,避免冗余重复
内容的提问来源于stack exchange,提问作者Nermin ÖZCAN
相关产品推荐
相关产品推荐

