You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

构建不平衡二分类模型时,如何排除标签列删除异常值行?

Fixing Outlier Removal for Imbalanced Binary Classification Datasets

Ah, I see the issue here! Your current code calculates z-scores across all columns in your DataFrame—including the label column. Since your label is a binary variable (0/1) with a massive imbalance (only 5% are class 0), the z-score for the label column will flag nearly all class 0 rows as outliers (they’re far from the mean value of ~0.95). That’s exactly why those critical minority samples are getting dropped accidentally.

The Correct Approach: Target Only Feature Columns

We need to isolate your feature columns (all columns except the label) for outlier detection, then apply the filter to the full DataFrame to keep the corresponding label values. Here’s how to implement this:

First, assume your label column is named target (replace this with your actual label column name):

import numpy as np
from scipy import stats

# Step 1: Identify feature columns (exclude the label column)
feature_cols = df.columns[df.columns != "target"]

# Step 2: Calculate z-scores ONLY for feature columns and create a filter mask
# We keep rows where ALL feature columns have z-scores < 3 in absolute value
outlier_mask = (np.abs(stats.zscore(df[feature_cols])) < 3).all(axis=1)

# Step 3: Apply the mask to clean your full DataFrame (features + label)
df_cleaned = df[outlier_mask]

Why This Works

By focusing solely on feature columns for z-score calculation, we avoid penalizing rows just because they belong to the minority class (0). The label column stays completely untouched during the outlier detection process, so your class 0 samples won’t be dropped unless their feature values are truly extreme.

Additional Tips for Imbalanced Datasets

Since your dataset is heavily imbalanced, you might want to be more cautious with outlier removal (class 0 samples are already scarce!). Here are two alternative approaches to consider:

  1. IQR Method (Less Sensitive to Extreme Values)
    The Interquartile Range (IQR) method is more robust for non-normally distributed data and less likely to filter out valid minority samples:

    # Calculate IQR for each feature column
    Q1 = df[feature_cols].quantile(0.25)
    Q3 = df[feature_cols].quantile(0.75)
    IQR = Q3 - Q1
    
    # Create mask to exclude rows where any feature is outside 1.5*IQR from Q1/Q3
    iqr_mask = ~((df[feature_cols] < (Q1 - 1.5 * IQR)) | (df[feature_cols] > (Q3 + 1.5 * IQR))).any(axis=1)
    df_cleaned_iqr = df[iqr_mask]
    
  2. Adjust Z-Score Threshold
    If you still prefer z-scores, try raising the threshold from 3 to 3.5 or 4. This reduces the strictness of outlier detection, which can help preserve more class 0 samples.

内容的提问来源于stack exchange,提问作者dododips

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 07:34:15