为何用np.nan替换异常值会清空列中所有非零数据?
异常值替换与填充问题排查
我需要将数据列中的异常值替换为np.nan(而非整列替换),之后用中位数填充所有np.nan值。目前尝试了两种方案:
- 第一种是
remove_outlier函数,它返回移除异常值行后的新数据集,从BsmtFinSF2变量的箱线图可见列中仍有数据,但该代码并非我编写。 - 自行编写的
drop_outliers函数出现异常:将列中所有数据替换为np.nan,导致该列仅剩余0或np.nan值,对应的BsmtFinSF2变量箱线图为空。
相关代码
他人编写的移除异常值代码
import numpy as np import pandas as pd def outlier_thresholds(dataframe, col_name, low_q = 0.05, up_q= 0.95): q1, q3 = np.nanpercentile(dataframe[col_name], [low_q, up_q]) iqr = q3 - q1 lower_thres = q1 - (1.5 * iqr) upper_thres = q3 + (1.5 * iqr) return lower_thres, upper_thres def check_outlier(dataframe, col_name): lower_thres, upper_thres = outlier_thresholds(dataframe, col_name) if dataframe[(dataframe[col_name] > upper_thres) | (dataframe[col_name] < lower_thres)].any(axis=None): return True return False outlier_list = [] for col in num_list: if check_outlier(df, col): outlier_list.append(col) def remove_outlier(dataframe, col_name): lower_thres, upper_thres = outlier_thresholds(dataframe, col_name) df_without = dataframe[~((dataframe[col_name] < lower_thres) | (dataframe[col_name] > upper_thres))] return df_without for col in outlier_list: new_df = remove_outlier(df, col)
自行编写的异常值替换为np.nan的代码
def drop_outliers(dataframe, col_name): lower_thres, upper_thres = outlier_thresholds(dataframe, col_name) dataframe.loc[(dataframe[col_name] < lower_thres) | (dataframe[col_name] > upper_thres), dataframe[col_name]] = np.nan for col in outlier_list: drop_outliers(df, col)
示例数据
sample_df = {0: 0.0, 1: 0.0, 2: 0.0, 3: 0.0, 4: 0.0, 5: 0.0, 6: 0.0, 7: 32.0, 8: 0.0, 9: 0.0, 10: 0.0, 11: 0.0, 12: 0.0, 13: 0.0, 14: 0.0, 15: 0.0, 16: 0.0, 17: 0.0, 18: 0.0, 19: 0.0, 20: 0.0, 21: 0.0, 22: 0.0, 23: 0.0, 24: 668.0, 25: 0.0, 26: 486.0, 27: 0.0, 28: 0.0, 29: 0.0}
问题排查与修正
核心问题
drop_outliers函数的loc索引写法错误:
# 错误写法 dataframe.loc[条件, dataframe[col_name]] = np.nan
这里的列索引使用dataframe[col_name]会触发整列匹配,导致所有行都被选中替换为np.nan。
修正后的代码
将列索引改为直接传入列名字符串:
def drop_outliers(dataframe, col_name): lower_thres, upper_thres = outlier_thresholds(dataframe, col_name) # 修正列索引为col_name字符串 dataframe.loc[(dataframe[col_name] < lower_thres) | (dataframe[col_name] > upper_thres), col_name] = np.nan for col in outlier_list: drop_outliers(df, col)
补充中位数填充代码
完成异常值替换后,可通过以下代码用中位数填充np.nan:
for col in outlier_list: median_val = df[col].median() df[col].fillna(median_val, inplace=True)
内容的提问来源于stack exchange,提问作者Zek
相关产品推荐
相关产品推荐

