构建不平衡二分类模型时,如何排除标签列删除异常值行?
Ah, I see the issue here! Your current code calculates z-scores across all columns in your DataFrame—including the label column. Since your label is a binary variable (0/1) with a massive imbalance (only 5% are class 0), the z-score for the label column will flag nearly all class 0 rows as outliers (they’re far from the mean value of ~0.95). That’s exactly why those critical minority samples are getting dropped accidentally.
The Correct Approach: Target Only Feature Columns
We need to isolate your feature columns (all columns except the label) for outlier detection, then apply the filter to the full DataFrame to keep the corresponding label values. Here’s how to implement this:
First, assume your label column is named target (replace this with your actual label column name):
import numpy as np from scipy import stats # Step 1: Identify feature columns (exclude the label column) feature_cols = df.columns[df.columns != "target"] # Step 2: Calculate z-scores ONLY for feature columns and create a filter mask # We keep rows where ALL feature columns have z-scores < 3 in absolute value outlier_mask = (np.abs(stats.zscore(df[feature_cols])) < 3).all(axis=1) # Step 3: Apply the mask to clean your full DataFrame (features + label) df_cleaned = df[outlier_mask]
Why This Works
By focusing solely on feature columns for z-score calculation, we avoid penalizing rows just because they belong to the minority class (0). The label column stays completely untouched during the outlier detection process, so your class 0 samples won’t be dropped unless their feature values are truly extreme.
Additional Tips for Imbalanced Datasets
Since your dataset is heavily imbalanced, you might want to be more cautious with outlier removal (class 0 samples are already scarce!). Here are two alternative approaches to consider:
IQR Method (Less Sensitive to Extreme Values)
The Interquartile Range (IQR) method is more robust for non-normally distributed data and less likely to filter out valid minority samples:# Calculate IQR for each feature column Q1 = df[feature_cols].quantile(0.25) Q3 = df[feature_cols].quantile(0.75) IQR = Q3 - Q1 # Create mask to exclude rows where any feature is outside 1.5*IQR from Q1/Q3 iqr_mask = ~((df[feature_cols] < (Q1 - 1.5 * IQR)) | (df[feature_cols] > (Q3 + 1.5 * IQR))).any(axis=1) df_cleaned_iqr = df[iqr_mask]Adjust Z-Score Threshold
If you still prefer z-scores, try raising the threshold from 3 to 3.5 or 4. This reduces the strictness of outlier detection, which can help preserve more class 0 samples.
内容的提问来源于stack exchange,提问作者dododips

