You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

深度学习文本分类任务中更换数据集导致模型精度波动差异大的原因探究

Why Your BiLSTM Model Has Wild Accuracy Fluctuations on Some Datasets

Let’s dig into why you’re seeing such inconsistent accuracy (from 50% to 97%) with certain datasets, while others give steady ~95% results. I’ll tie this directly to your code and common text classification pitfalls:

1. The test_model Function Is Causing Critical Data Leakage

Looking at your code, this function is a major red flag:

def test_model(model, epoch_stop):
    model.fit(X_test , Y_test , epochs=epoch_stop , batch_size=batch_size , verbose=0)
    results = model.evaluate(X_test, Y_test)
    return results

You’re training your model on the test set here! This completely contaminates your test results—when you run this, you’re teaching the model the exact answers it needs to predict for the test data, leading to artificially high accuracy. If you’re accidentally using this function for some datasets but not others, that’s an immediate source of wild fluctuations.

Fix this by removing the fit call entirely; testing should only involve evaluating a trained model on unseen data:

def test_model(model):
    results = model.evaluate(X_test, Y_test)
    return results

2. Dataset Quality & Distribution Is the Root Cause

The biggest driver of inconsistent results is almost always the data itself:

  • Class imbalance: If your volatile datasets have skewed class distributions (e.g., one class makes up 90% of the data), a model can hit 90% accuracy just by guessing the majority class. But if the train/test split puts most of the minority class in the test set, accuracy plummets to 50%. Your stable datasets likely have balanced, consistent class distributions.
  • Label noise & data consistency: Volatile datasets probably have more mislabeled samples, irrelevant text, or mix multiple unrelated domains (e.g., mixing medical reviews with tech support tickets). Models struggle to learn meaningful patterns from noisy, disjoint data, leading to unpredictable performance. Your stable datasets likely have clean, domain-consistent text.
  • Small sample sensitivity: Even if datasets are the same size, volatile ones might have small subsets of rare classes. A single bad train/test split can leave the model with no examples of a rare class in training, making it impossible to predict that class in testing.

3. Training Instabilities in Your Model Setup

Your current training pipeline has several factors that amplify fluctuations:

  • Too few epochs: You’re only training for 4 epochs. For simple datasets, this might be enough to converge to high accuracy, but for complex/noisy datasets, 4 epochs barely scratches the surface—your model is still learning random patterns instead of generalizable features. Add early stopping to let the model train until it stops improving:
    from tensorflow.keras.callbacks import EarlyStopping
    early_stop = EarlyStopping(monitor='val_loss', patience=3, restore_best_weights=True)
    history = lstm_model.fit(X_train, label_train, epochs=20, verbose=True, 
                             validation_data=(X_test, label_test), batch_size=64,
                             callbacks=[early_stop])
    
  • Model capacity & randomness: Your BiLSTM uses 64 units, which might be too small for volatile datasets. Smaller models are more sensitive to random initialization and noise—tiny changes in weight initialization can lead to huge swings in performance. Your stable datasets have clear enough patterns that even a small model can learn them consistently.
  • Inconsistent preprocessing (unseen in your code): You didn’t share your text preprocessing steps (tokenization, word embedding setup, cleaning). If you’re using different preprocessing for different datasets (e.g., random embeddings for some, pre-trained embeddings for others; skipping stopword removal for some), this will directly cause accuracy fluctuations.

4. Train-Test Split Limitations (Even with random_state=42)

While random_state=42 fixes the split seed, it doesn’t fix bad splits if the dataset is inherently messy. For example, if all samples of a rare class are clustered together in the dataset, a fixed split might still put all of them in the test set. For volatile datasets, use stratified splitting to preserve class distributions:

X_train, X_test, label_train, label_test = train_test_split(X, Y, test_size=0.2, 
                                                            random_state=42,
                                                            stratify=Y)

Quick Fixes to Stabilize Performance

  1. Delete the flawed test_model function and stop training on test data.
  2. Add stratified train/test splits for all datasets.
  3. Implement early stopping to let models train to convergence.
  4. Audit your volatile datasets for class imbalance, label noise, and domain consistency.
  5. Ensure all datasets use identical preprocessing pipelines.

内容的提问来源于stack exchange,提问作者eman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 17:33:11