如何选择train_test_split的RandomState?如何可靠报告模型准确率?
random_state Causes Fluctuations? Great question—this is a super common pitfall when documenting machine learning experiments, especially for academic reports where rigor matters a lot. Let’s break down your options and the best practices to get a reliable, representative performance measure:
Avoid using the highest observed accuracy
Picking the single best score from differentrandom_stateruns is a bad idea. That result is likely inflated by chance (you just got a lucky train-test split) and won’t reflect how your model actually performs on unseen data. It’s essentially cherry-picking, which will weaken the credibility of your report.Averaging scores is better, but pair it with variability metrics
Taking the average accuracy across multiplerandom_statesplits is a step in the right direction—it gives you a sense of the model’s typical performance. But don’t stop at just the average! You should also report the standard deviation (or standard error) of those scores. This tells your readers how consistent the model is; a high standard deviation means your model’s performance is highly sensitive to the specific train-test split, which is a key limitation to highlight.The most robust approach: K-Fold Cross-Validation
Instead of relying on multiple random train-test splits, use K-Fold Cross-Validation (10-fold is a standard choice). Here’s the gist:- Split your entire dataset into K equal, non-overlapping folds.
- Train your model K times: each time, use K-1 folds as training data and the remaining 1 fold as test data.
- Calculate the average of the K accuracy scores, plus the standard deviation.
This method uses all your data for both training and testing (in rotation), so it’s far more reliable than random splits. You should still set arandom_statefor the cross-validation split itself to ensure your results are reproducible.
What to include in your report
- Lead with the mean accuracy (from cross-validation or averaged random splits) as your primary metric.
- Add the standard deviation in parentheses or with a ± symbol, e.g., "87.2% ± 1.3%". This shows both typical performance and consistency.
- If you use random splits, specify how many splits you ran (e.g., "average of 20 random 80-20 train-test splits with unique
random_statevalues"). - Always note your split ratio, cross-validation parameters, and any
random_statevalues used for reproducibility.
内容的提问来源于stack exchange,提问作者Prabhjeet Singh

