You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何遍历多份DataFrame并为各特征绘制对比直方图

问题描述

我有两个DataFrame:df1和df2,数据如下:

# df1数据
Age BsHgt_M BsWgt_Kg    GOAT-MBOAT4_F_BM    TCF7L2_M_BM UCP2_M_BM
23.0    1.84    113.0   -1.623634   0.321379    0.199183
23.0    1.68    113.9   -1.073523   -0.957523   0.549469
24.0    1.60    86.4    -0.270883   -0.004106   1.479865
20.0    1.59    99.2    -0.218071   0.568458    -0.398410

# df2数据
Age BsHgt_M BsWgt_Kg    GOAT-MBOAT4_F_BM    TCF7L2_M_BM UCP2_M_BM
29.0    1.94    123.0   -1.623676   0.321379    0.199183
30.0    1.61    113.9   -1.073523   -0.957523   0.549469
44.0    1.30    56.4    -0.270883   -0.004106   1.479865
30.0    1.19    91.2    -0.218071   0.568458    -0.398410

我已经实现了遍历df1各列绘制单份直方图的代码:

import matplotlib.pyplot as plt

fig, axs = plt.subplots(len(df1.columns), figsize=(10,50))
for n, col in enumerate(df1.columns):
    df1[col].hist(ax=axs[n],legend=True)

现在需要遍历这两个DataFrame,为每个对应特征(比如df1['Age']和df2['Age'])在同一图表中绘制直方图,或者绘制同尺度的并排直方图,求实现方法。


方案1:同一子图叠加直方图(带透明度区分)

这种方法适合对比两个数据集同一特征的分布重叠情况,通过设置alpha参数(透明度)区分两个直方图:

import matplotlib.pyplot as plt

# 获取列数,确保两个DataFrame列名一致
num_cols = len(df1.columns)
fig, axs = plt.subplots(num_cols, figsize=(10, 5*num_cols))

for n, col in enumerate(df1.columns):
    # 绘制df1的直方图,设置透明度和标签
    df1[col].hist(ax=axs[n], alpha=0.5, label='df1', bins=10)
    # 绘制df2的直方图,同样设置透明度和标签
    df2[col].hist(ax=axs[n], alpha=0.5, label='df2', bins=10)
    # 设置子图标题、图例和坐标轴标签
    axs[n].set_title(f'Distribution of {col}')
    axs[n].legend()
    axs[n].set_xlabel(col)
    axs[n].set_ylabel('Frequency')

# 调整子图间距,避免标题和内容重叠
plt.tight_layout()
plt.show()

方案2:同尺度并排直方图

如果希望更清晰地分开对比,可以为每个特征创建两个并排子图,并且统一坐标轴尺度,保证对比公平性:

import matplotlib.pyplot as plt

num_cols = len(df1.columns)
# 创建num_cols行2列的子图布局
fig, axs = plt.subplots(num_cols, 2, figsize=(12, 5*num_cols))

for n, col in enumerate(df1.columns):
    # 绘制df1的直方图
    df1[col].hist(ax=axs[n, 0], bins=10)
    axs[n, 0].set_title(f'{col} - df1')
    axs[n, 0].set_xlabel(col)
    axs[n, 0].set_ylabel('Frequency')
    
    # 绘制df2的直方图
    df2[col].hist(ax=axs[n, 1], bins=10)
    axs[n, 1].set_title(f'{col} - df2')
    axs[n, 1].set_xlabel(col)
    axs[n, 1].set_ylabel('Frequency')
    
    # 统一x轴和y轴范围,确保尺度一致
    x_min = min(df1[col].min(), df2[col].min())
    x_max = max(df1[col].max(), df2[col].max())
    # 计算两个数据集的最大频数,统一y轴上限
    y_max = max(df1[col].value_counts(bins=10).max(), df2[col].value_counts(bins=10).max())
    
    axs[n, 0].set_xlim(x_min, x_max)
    axs[n, 1].set_xlim(x_min, x_max)
    axs[n, 0].set_ylim(0, y_max + 0.5)
    axs[n, 1].set_ylim(0, y_max + 0.5)

plt.tight_layout()
plt.show()

注意事项

  • 确保两个DataFrame的列名完全一致,否则对应特征会不匹配;
  • bins参数可以根据数据分布调整,让直方图展示更合理;
  • 方案2中统一坐标轴尺度是关键,避免因尺度差异导致的视觉误导。

内容的提问来源于stack exchange,提问作者paul raj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 22:18:20