Pandas多分类列基于value_counts()绘制网格条形图的简洁方法
需求说明
我需要为Pandas DataFrame中的所有分类列绘制唯一值计数条形图,效果和数值列调用df.hist()生成的直方图一致,具体要求:
- 优先使用面向对象的绘图方式,逻辑更自然、表述更明确
- 所有子图以网格形式排布在单个Figure中,布局效果参考
df.hist()的输出
目前我写的代码可以实现效果,但写法繁琐,需要手动调用Matplotlib创建Figure、手动移除最后一排未使用的空白子图。我注意到pandas.Series.plot提供了subplots和layout参数,理论上可以直接实现需求,但遍历列传入参数时始终无法得到正确结果,现寻求更简洁紧凑的实现方式。
现有可运行但繁琐的代码
import math import matplotlib.pyplot as plt # 定义子图网格维度 nr_of_plots = len(ames_train_categorical.columns) nr_of_plots_per_row = 4 nr_of_rows = math.ceil(nr_of_plots / nr_of_plots_per_row) # 创建Figure和子图对象 figure, axes = plt.subplots(nrows=nr_of_rows, ncols=nr_of_plots_per_row, figsize=(25, 50)) figure.subplots_adjust(hspace=0.5) # 逐列绘制条形图 i, j = 0, 0 for column_name in ames_train_categorical: if ames_train_categorical[column_name].nunique() <= 30: axes[i][j].set_title(column_name) ames_train_categorical[column_name].value_counts().plot(kind='bar', ax=axes[i][j]) j += 1 if j % nr_of_plots_per_row == 0: i += 1 j = 0 # 清理未使用的空白子图 axes_flattened = axes.flatten() for ax in axes_flattened: if not ax.has_data(): ax.remove()
备选方案的缺陷
使用pyplot状态机模式虽然代码量极少,但每个图表会单独生成一个Figure,无法实现规整的网格排布,代码如下:
for column_name in ames_train_categorical: ames_train_categorical[column_name].value_counts().plot(kind='bar') plt.show()
期望效果

简洁实现方案
不需要手动维护子图行列索引、手动移除空白子图,用pandas原生绘图接口即可实现和df.hist()一致的网格排布效果,核心是直接对DataFrame调用绘图方法,而非逐列遍历Series绘图。
推荐写法(适配性最强)
import matplotlib.pyplot as plt # 按原逻辑过滤唯一值数量≤30的分类列,避免分类过多导致条形图拥挤 target_cols = ames_train_categorical.columns[ames_train_categorical.nunique() <= 30] # 逐列计算值计数并绘图,pandas会自动创建适配的子图网格 axes = ames_train_categorical[target_cols].apply( lambda col: col.value_counts().plot(kind='bar') ).axes # 调整画布尺寸和子图间距,和原有参数保持一致 plt.gcf().set_size_inches(25, 5 * ((len(target_cols) + 3) // 4)) plt.subplots_adjust(hspace=0.5, wspace=0.3) # 关闭无内容的空白子图,无需remove for ax in axes.flat: if not ax.patches: # 条形图通过patches判断是否有绘制内容,比has_data更准确 ax.set_axis_off()
极简写法(代码最短)
如果不需要考虑不同列分类值不对齐的问题,可以直接传入布局参数,让pandas自动完成所有网格计算:
# 预计算所有待绘制列的值计数 count_df = ames_train_categorical[target_cols].apply(lambda s: s.value_counts()).fillna(0) axes = count_df.plot( kind='bar', subplots=True, layout=(-1, 4), # 固定每行4个子图,行数自动计算 figsize=(25, 50), legend=False ) plt.tight_layout(h_pad=3)
注意:极简写法会自动将所有分类值对齐到同一x轴刻度,如果各列的分类取值差异较大,会出现大量空白刻度位置,日常使用优先选推荐写法适配性更好。
内容的提问来源于stack exchange,提问作者luukburger
相关产品推荐
相关产品推荐

