You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何用np.nan替换异常值会清空列中所有非零数据?

异常值替换与填充问题排查

我需要将数据列中的异常值替换为np.nan(而非整列替换),之后用中位数填充所有np.nan值。目前尝试了两种方案:

  • 第一种是remove_outlier函数,它返回移除异常值行后的新数据集,从BsmtFinSF2变量的箱线图可见列中仍有数据,但该代码并非我编写。
  • 自行编写的drop_outliers函数出现异常:将列中所有数据替换为np.nan,导致该列仅剩余0或np.nan值,对应的BsmtFinSF2变量箱线图为空。

相关代码

他人编写的移除异常值代码

import numpy as np
import pandas as pd

def outlier_thresholds(dataframe, col_name, low_q = 0.05, up_q= 0.95):
    q1, q3 = np.nanpercentile(dataframe[col_name], [low_q, up_q])
    iqr = q3 - q1
    lower_thres = q1 - (1.5 * iqr)
    upper_thres = q3 + (1.5 * iqr)
    return lower_thres, upper_thres

def check_outlier(dataframe, col_name):
    lower_thres, upper_thres = outlier_thresholds(dataframe, col_name)
    if dataframe[(dataframe[col_name] > upper_thres) | 
                 (dataframe[col_name] < lower_thres)].any(axis=None):
        return True
    return False

outlier_list = []
for col in num_list:
    if check_outlier(df, col):
        outlier_list.append(col)

def remove_outlier(dataframe, col_name):
    lower_thres, upper_thres = outlier_thresholds(dataframe, col_name)
    df_without = dataframe[~((dataframe[col_name] < lower_thres) | (dataframe[col_name] > upper_thres))]
    return df_without

for col in outlier_list:
    new_df = remove_outlier(df, col)

自行编写的异常值替换为np.nan的代码

def drop_outliers(dataframe, col_name):
    lower_thres, upper_thres = outlier_thresholds(dataframe, col_name)
    dataframe.loc[(dataframe[col_name] < lower_thres) | (dataframe[col_name] > upper_thres), dataframe[col_name]] = np.nan

for col in outlier_list:
    drop_outliers(df, col)

示例数据

sample_df = {0: 0.0,
 1: 0.0,
 2: 0.0,
 3: 0.0,
 4: 0.0,
 5: 0.0,
 6: 0.0,
 7: 32.0,
 8: 0.0,
 9: 0.0,
 10: 0.0,
 11: 0.0,
 12: 0.0,
 13: 0.0,
 14: 0.0,
 15: 0.0,
 16: 0.0,
 17: 0.0,
 18: 0.0,
 19: 0.0,
 20: 0.0,
 21: 0.0,
 22: 0.0,
 23: 0.0,
 24: 668.0,
 25: 0.0,
 26: 486.0,
 27: 0.0,
 28: 0.0,
 29: 0.0}

问题排查与修正

核心问题

drop_outliers函数的loc索引写法错误:

# 错误写法
dataframe.loc[条件, dataframe[col_name]] = np.nan

这里的列索引使用dataframe[col_name]会触发整列匹配,导致所有行都被选中替换为np.nan。

修正后的代码

将列索引改为直接传入列名字符串:

def drop_outliers(dataframe, col_name):
    lower_thres, upper_thres = outlier_thresholds(dataframe, col_name)
    # 修正列索引为col_name字符串
    dataframe.loc[(dataframe[col_name] < lower_thres) | (dataframe[col_name] > upper_thres), col_name] = np.nan

for col in outlier_list:
    drop_outliers(df, col)

补充中位数填充代码

完成异常值替换后,可通过以下代码用中位数填充np.nan:

for col in outlier_list:
    median_val = df[col].median()
    df[col].fillna(median_val, inplace=True)

内容的提问来源于stack exchange,提问作者Zek

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 18:23:14