IQR法移除异常值无效:DataFrame维度未发生改变
问题分析与解决方案
核心错误点
单列操作无法删除整行
你执行的df[c].drop(index=upper_array,inplace=True)是对单个列(Series对象)进行删除操作,只会将该列对应位置的值设为NaN,不会删除整个DataFrame的行,因此DataFrame的总维度不会发生变化。索引处理逻辑错误
np.where(df[c]>=upper)[0]返回的是符合条件的行位置数组(如array([10, 25, 32])),而非布尔掩码。直接使用df= df[~upper_array]是错误的——~是对布尔值取反,对整数数组取反会生成负数,无法作为有效的索引条件。循环中逐列删除的风险
即使你修改为删除整行,循环中逐列删除会导致后续列的行位置发生偏移(因为前面已经删除过行),但你仍使用原始的行位置来定位异常值,会导致遗漏或错误删除。
正确实现方式
方案1:收集所有异常值索引后一次性删除
这种方式会收集两列中所有异常值对应的行索引,去重后一次性删除,避免循环中的索引偏移问题:
import numpy as np import pandas as pd # 用集合存储异常值索引,自动去重 outlier_indices = set() # 遍历目标列 for col in df.columns[:2]: Q1 = np.percentile(df[col], 25) Q3 = np.percentile(df[col], 75) IQR = Q3 - Q1 upper_bound = Q3 + 1.5 * IQR lower_bound = Q1 - 1.5 * IQR # 获取当前列的异常值行索引 current_outliers = df[(df[col] >= upper_bound) | (df[col] <= lower_bound)].index outlier_indices.update(current_outliers) # 删除所有异常值行 df = df.drop(index=outlier_indices)
方案2:用布尔掩码过滤行
通过生成布尔掩码,逐步筛选出两列都无异常的行:
import numpy as np import pandas as pd # 初始化掩码:默认保留所有行 keep_mask = pd.Series([True] * len(df), index=df.index) for col in df.columns[:2]: Q1 = np.percentile(df[col], 25) Q3 = np.percentile(df[col], 75) IQR = Q3 - Q1 upper_bound = Q3 + 1.5 * IQR lower_bound = Q1 - 1.5 * IQR # 更新掩码:仅保留当前列非异常的行 keep_mask &= (df[col] > lower_bound) & (df[col] < upper_bound) # 应用掩码过滤DataFrame df = df[keep_mask]
验证效果
执行完上述代码后,你可以通过df.shape查看新的维度,确认异常值行已被删除。
内容的提问来源于stack exchange,提问作者fffff
相关产品推荐
相关产品推荐

