You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

稀疏标签下深度逻辑回归应用:不平衡数据集训练问题问询

Alright, let's break down your problem step by step—your 99:1 class imbalance is the elephant in the room here, and pairing that with MSE loss is creating some hidden issues you might not be seeing at first glance.

Core Problem Analysis

1. MSE Loss Is Terrible for Imbalanced Data

  • You're using mean_squared_error as your loss function, but this metric naturally favors the majority class (negative samples, 99% of your data). Think about it: if your model just predicts a value close to 0 for every sample, the MSE will be extremely low (since 99k negative samples will contribute almost nothing to the loss). Even if it completely misses all 1k positive samples, the overall loss will still look like it's "improving"—but that improvement is just the model optimizing for the easy, abundant negative samples, not learning to detect positives.
  • For example: Predicting 0.01 for all samples gives an MSE of ~0.001, which looks great on paper, but your positive sample predictions are worthless.

2. Model & Training Dynamics Are Working Against You

  • Your 4-layer DNN with ReLU activations is fine structurally, but imbalanced data can easily trigger "dead ReLU neurons"—if the model keeps outputting low values (to fit negatives), many ReLU units will stay stuck at 0, making their weights impossible to update.
  • The sigmoid output layer amplifies this issue: when predictions are close to 0, the sigmoid's gradient is almost flat. Adam optimizer will struggle to push the model to learn positive sample patterns because the gradient signal from positives is drowned out by the massive number of negatives.
Practical Fixes to Try

1. Swap Loss Function (Highest Priority)

  • Ditch MSE immediately and use binary_crossentropy—this is the standard loss for binary classification, and it penalizes misclassifications of rare positives more heavily than MSE does.
  • If binary crossentropy still isn't enough, use weighted crossentropy: assign a higher weight to positive samples (e.g., a weight of 99, matching the inverse of their class ratio). This tells the model that misclassifying a positive is as bad as misclassifying 99 negatives. Here's how you might implement this in code (using TensorFlow/Keras):
    # Calculate class weights
    pos_weight = (1 / 1000) * (100000 / 2)
    neg_weight = (1 / 99000) * (100000 / 2)
    class_weights = {0: neg_weight, 1: pos_weight}
    
    # Pass to fit()
    model.fit(X_train, y_train, class_weight=class_weights, ...)
    
  • For even better results, try Focal Loss—it reduces the loss contribution of easy-to-classify negative samples, forcing the model to focus on the hard positive cases that matter most.

2. Fix the Dataset Imbalance

  • Oversample positive samples: Duplicate your 1k positive samples to balance the ratio, or use SMOTE to generate synthetic positive samples (just be careful not to overfit to fake data).
  • Undersample negative samples: Randomly remove some negative samples to get a more balanced ratio (like 1:10 positive-to-negative). Combine this with cross-validation to avoid losing critical negative sample patterns.
  • Stratified batching: Ensure every training batch has a meaningful number of positive samples (e.g., 10-20% positives). This way, the model gets regular signal from positives instead of going through dozens of batches with zero positive examples.

3. Adjust Training & Evaluation

  • Stop judging success by loss alone: MSE (or even crossentropy loss) doesn't tell the whole story. Track metrics that matter for imbalanced data: Recall (how many positives you actually catch), Precision (how many predicted positives are real), F1-Score, and AUC-ROC. These metrics will show you if your model is actually learning to detect positives, not just fitting negatives.
  • Use He initialization: For ReLU layers, He initialization helps prevent dead neurons by setting weights to a scale that keeps activations in a healthy range. In Keras, just set kernel_initializer='he_normal' for your Dense layers.
  • Early stopping with validation metrics: Use a validation set to monitor AUC or F1-score, and stop training when these metrics stop improving. This prevents overfitting to the majority class.

4. Tweak Your Model Structure

  • Add Dropout layers (rate 0.2-0.5) after ReLU activations to reduce overfitting to the abundant negative samples.
  • Add a BatchNormalization layer right before the sigmoid output. This stabilizes the values fed into sigmoid, preventing gradient vanishing and making it easier for the model to adjust predictions for positives.

内容的提问来源于stack exchange,提问作者Vincent Teyssier

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:58:01