基于四分位距筛选数据范围内有效值的Python代码错误修正问询
Fixing Your IQR Outlier Removal Code
Let's break down what's going wrong with your code and fix it step by step:
What's Causing the Issues?
- Incorrect Logical Operator: You used
|(OR) instead of&(AND). The condition(df[column] > Low ) | (df[column] < High)will always beTruefor every value—any number is either greater than the lower bound or less than the upper bound. This means no rows get filtered out at all. You need to check that values are both above the lower bound and below the upper bound. - Breaking the DataFrame Structure: When you slice with
[column]at the end of the line, you're converting your DataFrame into a single-column Series. By the next iteration of the loop, there are no other columns left to process, which causes the error you're seeing. - Updating Data Mid-Loop (Optional): If you filter the DataFrame in each loop, you'll be calculating IQR based on already filtered data for subsequent columns. This is rarely intended—most of the time, you want to use the original dataset's quantiles to set thresholds.
Corrected Code (Common Use Case)
This version keeps rows where all columns fall within their respective IQR ranges, using the original data's quantiles to set thresholds:
import numpy as np import pandas as pd # Start with a copy of your original data filtered_df = data.copy() # Initialize a mask to keep track of valid rows (starts as all True) valid_rows = pd.Series([True] * len(filtered_df), index=filtered_df.index) for column in data.columns: # Calculate IQR bounds using the ORIGINAL data, not filtered data Q1 = np.quantile(data[column], 0.25) Q3 = np.quantile(data[column], 0.75) IQR = Q3 - Q1 lower_bound = Q1 - 3 * IQR upper_bound = Q3 + 3 * IQR # Update the mask: only keep rows where this column is within bounds valid_rows &= (filtered_df[column] >= lower_bound) & (filtered_df[column] <= upper_bound) # Apply the mask to get your final cleaned DataFrame filtered_df = filtered_df[valid_rows]
Alternative: Filter Columns One by One
If you want to filter rows for each column sequentially (using the already filtered data to calculate IQR for the next column), use this version:
import numpy as np import pandas as pd filtered_df = data.copy() for column in data.columns: # Calculate bounds using the current filtered data Q1 = np.quantile(filtered_df[column], 0.25) Q3 = np.quantile(filtered_df[column], 0.75) IQR = Q3 - Q1 lower_bound = Q1 - 3 * IQR upper_bound = Q3 + 3 * IQR # Filter rows for this column, keep all columns intact filtered_df = filtered_df[(filtered_df[column] >= lower_bound) & (filtered_df[column] <= upper_bound)]
Quick Note on IQR Multiplier
You're using 3 * IQR to set bounds, which targets extreme outliers. The standard for mild outliers is 1.5 * IQR—adjust this number based on how strict you want your outlier removal to be.
内容的提问来源于stack exchange,提问作者Suraj Goswami
相关产品推荐
相关产品推荐

