如何消除Matplotlib分组柱状图中条形间的多余间隙
解决分组柱状图中Neutral类别条形间隙问题
问题原因
你遇到的间隙是因为label_percentages中Neutral类别对应的部分预测类(preds)值为0,pandas的plot.bar会为所有存在的preds列预留位置,零值的条形被隐藏但组内的间距仍被保留,导致出现空隙。另外align='center'和pandas分组柱状图的默认布局逻辑冲突,加剧了间隙问题。
解决方案
方法1:过滤全零列,只保留有数据的预测类
先移除所有在整个DataFrame中全为0的preds列,确保每个类别下的预测类都是有数据的,这样pandas绘制时不会预留空位置:
label_counts = df.groupby(['labels', 'preds']).size().unstack(fill_value=0) # Calculate percentages label_percentages = label_counts.div(label_counts.sum(axis=1), axis=0) * 100 label_percentages = label_percentages.loc[(label_percentages != 0).any(axis=1)] # 新增:过滤掉全零的预测类列 label_percentages = label_percentages.drop(columns=label_percentages.columns[label_percentages.sum() == 0]) # Plotting plt.figure(figsize=(10, 6), dpi=400) colors = ['coral', 'wheat', 'y'] # 去掉align='center',使用默认布局 label_percentages.plot(kind='bar', width=0.6, figsize=(10, 6), edgecolor='black', color=colors) plt.title('Distribution of Incorrect Predictions', fontweight='bold') plt.xlabel('Correct Label', fontweight='bold') plt.ylabel('% of incorrect predictions', fontweight='bold') plt.legend(title='Predicted Class', title_fontsize='medium') plt.tight_layout() plt.show()
方法2:使用Matplotlib原生API手动绘制(精确控制)
如果方法1无效,直接用Matplotlib原生代码绘制,完全掌控每个条形的位置和宽度,避免pandas自动布局的问题:
import numpy as np label_counts = df.groupby(['labels', 'preds']).size().unstack(fill_value=0) label_percentages = label_counts.div(label_counts.sum(axis=1), axis=0) * 100 label_percentages = label_percentages.loc[(label_percentages != 0).any(axis=1)] # 过滤全零列 label_percentages = label_percentages.drop(columns=label_percentages.columns[label_percentages.sum() == 0]) # 准备绘图数据 labels = label_percentages.index pred_classes = label_percentages.columns num_labels = len(labels) num_preds = len(pred_classes) bar_width = 0.8 / num_preds # 组内条形总宽度占满刻度位置 x = np.arange(num_labels) plt.figure(figsize=(10, 6), dpi=400) # 逐个绘制每个预测类的条形 for i, pred in enumerate(pred_classes): plt.bar(x + i * bar_width, label_percentages[pred], width=bar_width, edgecolor='black', color=colors[i]) # 设置x轴刻度和标签 plt.xticks(x + bar_width * (num_preds - 1) / 2, labels) plt.title('Distribution of Incorrect Predictions', fontweight='bold') plt.xlabel('Correct Label', fontweight='bold') plt.ylabel('% of incorrect predictions', fontweight='bold') plt.legend(title='Predicted Class', title_fontsize='medium') plt.tight_layout() plt.show()
关键说明
- 过滤全零列是核心:确保每个分组内的条形都是有数据的,不会有空位置预留。
- 避免使用
align='center':pandas的分组柱状图默认布局已经处理了组内对齐,手动设置反而会导致错位。 - 原生API绘制更灵活:当自动布局出现问题时,手动控制条形位置是最可靠的解决方式。
内容的提问来源于stack exchange,提问作者AnonymousMe
相关产品推荐
相关产品推荐

