如何基于两列字段将pandas数据的直方图拆分为多个子图
问题原因与解决方案
你之前出现KeyError是因为DataFrame取多列需要传入列表,正确写法是df[['id', 'barcode']],但即使修正该写法,直接使用hist的by参数也很难精准匹配5行2列的子图布局。下面提供两种可直接运行的实现方案:
方案1:pandas+matplotlib原生实现(无额外依赖)
该方案逻辑可控,自定义程度高,适配百万级数据无额外性能损耗。
import pandas as pd import matplotlib.pyplot as plt # 提前将accuracy转为数值型,避免字符串类型导致绘图异常 df['accuracy'] = pd.to_numeric(df['accuracy']) # 获取唯一id、barcode列表,排序保证子图位置固定 ids = sorted(df['id'].unique()) barcodes = sorted(df['barcode'].unique()) # 创建5行2列的子图,对应5种id、2种barcode fig, axes = plt.subplots(nrows=5, ncols=2, sharey='none', sharex='none', figsize=(20,20)) # 遍历双维度分组,对应到子图绘制直方图 for (id_val, barcode_val), group_df in df.groupby(['id', 'barcode']): row_idx = ids.index(id_val) col_idx = barcodes.index(barcode_val) current_ax = axes[row_idx, col_idx] # 可自行调整bins、颜色等参数 group_df['accuracy'].hist(ax=current_ax, bins=20, edgecolor='black') # 添加子图标识 current_ax.set_title(f"id: {id_val} | barcode: {barcode_val}") current_ax.set_xlabel("准确率") current_ax.set_ylabel("样本量") # 调整子图间距避免标题、标签重叠 plt.tight_layout() plt.show()
方案2:基于seaborn实现(代码更简洁)
如果接受引入第三方可视化库,用seaborn的FacetGrid可以更快捷的实现多维度分面子图:
import pandas as pd import matplotlib.pyplot as plt import seaborn as sns df['accuracy'] = pd.to_numeric(df['accuracy']) # 按id拆分行、barcode拆分列,自动生成5行2列的子图网格 g = sns.FacetGrid(df, row='id', col='barcode', height=4, aspect=1.5) # 每个子图绘制准确率直方图 g.map(plt.hist, 'accuracy', bins=20, edgecolor='black') # 统一设置轴标签和子图标题 g.set_axis_labels("准确率", "样本量") g.set_titles("id: {row_name} | barcode: {col_name}") plt.tight_layout() plt.show()
以上两种方案都可以满足你的需求,不需要使用其他更复杂的库。
内容的提问来源于stack exchange,提问作者Luke Roche
相关产品推荐
相关产品推荐

