如何遍历多份DataFrame并为各特征绘制对比直方图
问题描述
我有两个DataFrame:df1和df2,数据如下:
# df1数据 Age BsHgt_M BsWgt_Kg GOAT-MBOAT4_F_BM TCF7L2_M_BM UCP2_M_BM 23.0 1.84 113.0 -1.623634 0.321379 0.199183 23.0 1.68 113.9 -1.073523 -0.957523 0.549469 24.0 1.60 86.4 -0.270883 -0.004106 1.479865 20.0 1.59 99.2 -0.218071 0.568458 -0.398410 # df2数据 Age BsHgt_M BsWgt_Kg GOAT-MBOAT4_F_BM TCF7L2_M_BM UCP2_M_BM 29.0 1.94 123.0 -1.623676 0.321379 0.199183 30.0 1.61 113.9 -1.073523 -0.957523 0.549469 44.0 1.30 56.4 -0.270883 -0.004106 1.479865 30.0 1.19 91.2 -0.218071 0.568458 -0.398410
我已经实现了遍历df1各列绘制单份直方图的代码:
import matplotlib.pyplot as plt fig, axs = plt.subplots(len(df1.columns), figsize=(10,50)) for n, col in enumerate(df1.columns): df1[col].hist(ax=axs[n],legend=True)
现在需要遍历这两个DataFrame,为每个对应特征(比如df1['Age']和df2['Age'])在同一图表中绘制直方图,或者绘制同尺度的并排直方图,求实现方法。
方案1:同一子图叠加直方图(带透明度区分)
这种方法适合对比两个数据集同一特征的分布重叠情况,通过设置alpha参数(透明度)区分两个直方图:
import matplotlib.pyplot as plt # 获取列数,确保两个DataFrame列名一致 num_cols = len(df1.columns) fig, axs = plt.subplots(num_cols, figsize=(10, 5*num_cols)) for n, col in enumerate(df1.columns): # 绘制df1的直方图,设置透明度和标签 df1[col].hist(ax=axs[n], alpha=0.5, label='df1', bins=10) # 绘制df2的直方图,同样设置透明度和标签 df2[col].hist(ax=axs[n], alpha=0.5, label='df2', bins=10) # 设置子图标题、图例和坐标轴标签 axs[n].set_title(f'Distribution of {col}') axs[n].legend() axs[n].set_xlabel(col) axs[n].set_ylabel('Frequency') # 调整子图间距,避免标题和内容重叠 plt.tight_layout() plt.show()
方案2:同尺度并排直方图
如果希望更清晰地分开对比,可以为每个特征创建两个并排子图,并且统一坐标轴尺度,保证对比公平性:
import matplotlib.pyplot as plt num_cols = len(df1.columns) # 创建num_cols行2列的子图布局 fig, axs = plt.subplots(num_cols, 2, figsize=(12, 5*num_cols)) for n, col in enumerate(df1.columns): # 绘制df1的直方图 df1[col].hist(ax=axs[n, 0], bins=10) axs[n, 0].set_title(f'{col} - df1') axs[n, 0].set_xlabel(col) axs[n, 0].set_ylabel('Frequency') # 绘制df2的直方图 df2[col].hist(ax=axs[n, 1], bins=10) axs[n, 1].set_title(f'{col} - df2') axs[n, 1].set_xlabel(col) axs[n, 1].set_ylabel('Frequency') # 统一x轴和y轴范围,确保尺度一致 x_min = min(df1[col].min(), df2[col].min()) x_max = max(df1[col].max(), df2[col].max()) # 计算两个数据集的最大频数,统一y轴上限 y_max = max(df1[col].value_counts(bins=10).max(), df2[col].value_counts(bins=10).max()) axs[n, 0].set_xlim(x_min, x_max) axs[n, 1].set_xlim(x_min, x_max) axs[n, 0].set_ylim(0, y_max + 0.5) axs[n, 1].set_ylim(0, y_max + 0.5) plt.tight_layout() plt.show()
注意事项
- 确保两个DataFrame的列名完全一致,否则对应特征会不匹配;
bins参数可以根据数据分布调整,让直方图展示更合理;- 方案2中统一坐标轴尺度是关键,避免因尺度差异导致的视觉误导。
内容的提问来源于stack exchange,提问作者paul raj
相关产品推荐
相关产品推荐

