Python sklearn中train_test_split的test_size与10折交叉验证的关系咨询
交叉验证实现问题解答
- 你对
train_test_split代码的判断完全正确:这段代码仅会将数据集按70%训练、30%测试的比例做单次随机划分,不属于10折交叉验证的实现逻辑,无法满足你的10折验证需求。 - 你给出的KFold调用代码存在明显参数错误:参数
n_splits=2代表执行2折交叉验证,若要实现10折交叉验证,需要将该参数修改为n_splits=10。 - 适配你的二分类场景的优化建议:你要求每个折内两类样本占比均衡,更推荐使用分层10折交叉验证接口
StratifiedKFold,它会自动保证每一个折的类别分布和原始数据集一致,完全匹配你提到的「每个折每个类别各有约50个样本」的划分要求,参考实现代码如下:
from sklearn.model_selection import StratifiedKFold # 初始化分层10折交叉验证器,开启打乱保证划分随机性 skf = StratifiedKFold(n_splits=10, random_state=42, shuffle=True) # 遍历每个折进行模型训练评估 for train_index, test_index in skf.split(X, yy): x_train, x_test = X[train_index], X[test_index] y_train, y_test = yy[train_index], yy[test_index] # 下方写入你的模型训练、指标计算逻辑
内容的提问来源于stack exchange,提问作者Joe
相关产品推荐
相关产品推荐

