为何cross_val_score与train_test_split的RandomForestClassifier准确率不同?
为什么train_test_split的准确率比交叉验证高?
我用sklearn的RandomForestClassifier测试make_circles数据集,用train_test_split划分数据时准确率是0.89,用同参数的分类器做交叉验证时准确率约0.83,这是为什么?
代码示例
from sklearn.model_selection import cross_val_score, StratifiedKFold,GridSearchCV,train_test_split from sklearn.metrics import accuracy_score,f1_score,make_scorer from sklearn.ensemble import RandomForestClassifier from sklearn.datasets import make_circles import numpy as np np.random.seed(42) # 创建数据集: x, y = make_circles(n_samples=500, factor=0.1, noise=0.35, random_state=42) # 初始化分层划分: skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42) # 创建分类器: clf = RandomForestClassifier(random_state=42, max_depth=12,n_jobs=-1, oob_score=True,n_estimators=100,min_samples_leaf=10) # 交叉验证的平均准确率: results = np.mean(cross_val_score(clf, x, y, cv=skf,scoring=make_scorer(accuracy_score))) print("ACCURACY WITH CV = ",results)# 输出0.832 # 使用train_test_split划分数据 xtrain, xtest, ytrain, ytest = train_test_split(x, y, test_size=0.2) clf=RandomForestClassifier(random_state=42, max_depth=12,n_jobs=-1, oob_score=True,n_estimators=100,min_samples_leaf=10) clf.fit(xtrain,ytrain) ypred=clf.predict(xtest) print("ACCURACY WITHOUT CV = ",accuracy_score(ytest,ypred))# 输出0.89
运行结果
- 交叉验证平均准确率:0.83
- 单次划分测试准确率:0.89
原因解析
单次划分的随机“好运”
train_test_split默认未固定随机种子,这次划分的测试集恰好包含了更多模型容易预测的样本。make_circles添加噪声后,部分样本的类别边界更清晰,刚好被分到测试集,导致单次准确率偏高。交叉验证结果更可靠
5折交叉验证遍历了所有样本作为测试集的情况,平均结果消除了单次划分的随机性,更能反映模型在整个数据集上的真实泛化能力,是模型性能的可信估计。验证方法:固定划分种子
把train_test_split改成train_test_split(x, y, test_size=0.2, random_state=42)后再运行,准确率会接近交叉验证的0.83。固定种子后,划分逻辑和StratifiedKFold保持一致,避免了随机运气的影响。模型稳定性补充
你已经为随机森林设置了random_state=42,排除了模型自身的随机性干扰,所以差异完全来自数据划分的不同。
内容的提问来源于stack exchange,提问作者Андрей Малюк
相关产品推荐
相关产品推荐

