You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

课程项目中模型训练准确率100%、测试准确率60%-70%的问题咨询

Troubleshooting High Training Accuracy but Low Test Accuracy on Online News Popularity Dataset

Hey there, this classic gap between near-perfect training accuracy (≈100%) and mediocre test accuracy (60-70%) is almost always a sign of overfitting—your models are memorizing noise and idiosyncrasies in the training data instead of learning generalizable patterns that apply to unseen test data. Let’s break down the likely causes and actionable fixes:

1. First, Rule Out Dataset & Validation Issues

These are often the root cause, so start here:

  • Check for Data Leakage: Did you accidentally include test set information in your training pipeline? For example:
    • If you normalized/standardized features using statistics (mean, std) from the entire dataset (train + test) instead of only the training set, you’re leaking test data patterns into training. Always fit preprocessors on training data only, then transform test data with those same preprocessors.
    • Double-check your cross-validation setup: Are you using stratified splits? For classification tasks, StratifiedKFold ensures each fold has the same class distribution as the full dataset, which prevents skewed validation results.
  • Verify Train/Test Distribution: Are your training and test sets drawn from the same underlying distribution? For example, if your training set is dominated by news articles from a specific category/time period that’s underrepresented in the test set, models will fail to generalize. Use class distribution counts and feature distribution plots (like histograms of key features) to compare splits.
  • Assess Feature-to-Sample Ratio: If you have far more features than training samples, you’re facing the "curse of dimensionality"—models can easily memorize individual samples without learning meaningful patterns. Count your features vs. training samples; if features outnumber samples by a large margin, feature reduction is critical.

2. Tune Models to Reduce Overfitting

For SVMs:

  • Adjust Regularization Strength: The C parameter controls regularization—smaller values mean stronger regularization (penalizes complex models). Try reducing C (e.g., from default 1.0 to 0.1, 0.01) to force the model to prioritize simpler decision boundaries.
  • Switch to a Simpler Kernel: RBF kernels are powerful but prone to overfitting on noisy datasets. Try a linear kernel (kernel='linear') first—often, linear models generalize better even if they have slightly lower training accuracy.
  • Add Class Weights: If your dataset is imbalanced (e.g., most articles are "unpopular"), use class_weight='balanced' to prevent the model from just predicting the majority class.

For Random Forests:

  • Limit Tree Complexity: Reduce max_depth to prevent trees from growing too deep and memorizing individual samples. Start with a small depth (e.g., 5-10) and incrementally increase if validation accuracy improves.
  • Restrict Feature Access: Use max_features (e.g., 'sqrt' or a fixed number like 10) to limit how many features each tree can consider—this reduces correlation between trees and improves generalization.
  • Increase Minimum Sample Thresholds: Raise min_samples_split (minimum samples required to split a node) and min_samples_leaf (minimum samples required at a leaf node) to prevent trees from splitting on noisy, rare patterns.

3. Optimize Feature Engineering

  • Remove Redundant/Irrelevant Features:
    • Drop features with near-zero variance (they add no predictive value).
    • Use feature importance scores from your Random Forest (model.feature_importances_) to keep only the top 20-30 most impactful features.
    • Apply statistical tests (like chi-squared for categorical features, mutual information for continuous) to filter features that have no significant relationship with the target.
  • Apply Dimensionality Reduction: If you have many features, use PCA to compress them into a smaller set of orthogonal components that capture most of the variance. This reduces noise and the risk of overfitting.

4. Validate with a Baseline Model

Before diving deeper into complex models, train a simple baseline like Logistic Regression (with regularization via C). If the baseline has a smaller gap between training and test accuracy, that confirms your more complex models (SVM, RF) are overfitting. Use the baseline’s performance as a reference point for tuning more complex models.


内容的提问来源于stack exchange,提问作者Tulika Asthana

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:17:51