You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python plt.hist()绘制异常?DataFrame子集直方图显示不符求助

直方图异常问题:子集DataFrame出现原数据集不存在的数值

我有两个DataFrame,分别是merged和initial,其中initial是merged的子集。为对比两者数据分布,我绘制了两数据集各列的直方图,但发现initial的直方图中出现了本不该存在的数值(按子集定义,所有数值都应来自merged)。以fragC列为例:

  • merged的数值:[13.01, 46.03, 12.05, 64.08, 14.04]
  • initial的数值:[13.01, 64.08]
    但绘制出的直方图里,代表initial的“OPERA”分布存在异常。

我使用的代码

for column in common_columns:
    # Exclude the excluded_columns from the comparison
    if column not in excluded_columns:
        print("")
        our_values = df1[column].values
        opera_values = df2[column].values
        print(column)
        print(our_values)
        print(opera_values)
        # Plot the distribution for df1 and df2
        plt.figure(figsize=(10, 6))
        plt.hist(df1[column], bins=20, alpha=0.5, label='our dataset')
        plt.hist(df2[column], bins=20, alpha=0.5, label='OPERA')
        plt.xlabel('Values')
        plt.ylabel('Frequency')
        plt.title(f'Distribution Comparison for Column: {column}')
        plt.legend()
        plt.tight_layout()
        plt.show()

fragC列具体数据

merged的fragC列:{0: 13.01, 1: 46.03, 2: 12.05, 3: 64.08, 4: 14.04}
initial的fragC列:{0: 13.01, 1: 64.08}

可能的原因及解决办法

  1. bins自动划分导致视觉误导
    你设置了bins=20,但两个数据集的数值范围很小,过细的区间划分会让空bin或相邻bin被误判为异常值。解决方案是统一两个直方图的bins范围:

    import numpy as np
    
    for column in common_columns:
        if column not in excluded_columns:
            min_val = df1[column].min()
            max_val = df1[column].max()
            # 基于merged的数值范围生成统一的bins,减少数量避免过细划分
            bins = np.linspace(min_val, max_val, 10)
            
            plt.figure(figsize=(10, 6))
            plt.hist(df1[column], bins=bins, alpha=0.5, label='our dataset')
            plt.hist(df2[column], bins=bins, alpha=0.5, label='OPERA')
            plt.xlabel('Values')
            plt.ylabel('Frequency')
            plt.title(f'Distribution Comparison for Column: {column}')
            plt.legend()
            plt.tight_layout()
            plt.show()
    
  2. 隐形异常值或缺失值
    打印的数值可能不完整,需排查initial列是否存在不在merged中的数值或缺失值:

    for column in common_columns:
        if column not in excluded_columns:
            unique_initial = df2[column].unique()
            unique_merged = df1[column].unique()
            # 找出initial中不在merged里的数值
            unexpected_vals = [val for val in unique_initial if val not in unique_merged]
            print(f"{column}列的异常值:{unexpected_vals}")
            # 检查缺失值数量
            print(f"{column}列的缺失值数量:{df2[column].isna().sum()}")
    
  3. 浮点数精度误差
    看似相同的浮点数可能存在微小精度差异,导致被分到不同bins。可以对数值做四舍五入后再绘制:

    plt.hist(df1[column].round(2), bins=bins, alpha=0.5, label='our dataset')
    plt.hist(df2[column].round(2), bins=bins, alpha=0.5, label='OPERA')
    

内容的提问来源于stack exchange,提问作者C.D.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 11:45:25