如何替换train_test_split验证采样,转为全交叉验证方案?
改用全交叉验证的实现方案
原来的固定拆分方式会占用10%的数据作为独立验证集,无法参与模型训练。改用分层k折交叉验证后,所有非测试集的数据都会轮流作为训练集和验证集,最大化训练数据量,同时保持验证评估的可靠性。
代码实现
from sklearn.model_selection import StratifiedKFold, train_test_split # 提取特征和标签 X = data['text_with_tokeniz_lemmatiz'] y = data['toxic'] # 先拆分出10%的独立测试集(仅用于最终评估) X_train_val, X_test, y_train_val, y_test = train_test_split( X, y, test_size=0.1, shuffle=True, random_state=12345, stratify=y ) # 初始化9折分层交叉验证(保证每折类别分布与整体一致) skf = StratifiedKFold(n_splits=9, shuffle=True, random_state=12345) # 遍历每折进行训练与验证 for fold_idx, (train_indices, val_indices) in enumerate(skf.split(X_train_val, y_train_val)): # 划分当前折的训练集和验证集 X_train, X_val = X_train_val.iloc[train_indices], X_train_val.iloc[val_indices] y_train, y_val = y_train_val.iloc[train_indices], y_train_val.iloc[val_indices] # 替换为你的模型训练代码 # model = 你的模型实例 # model.fit(X_train, y_train) # val_performance = model.evaluate(X_val, y_val) print(f"第{fold_idx+1}折,验证集性能:{val_performance}") # 用全部训练验证数据训练最终模型,再评估测试集 final_model.fit(X_train_val, y_train_val) test_performance = final_model.evaluate(X_test, y_test) print(f"最终测试集性能:{test_performance}")
关键说明
- 选择
StratifiedKFold而非普通KFold:因为你的任务是二分类(toxic标签),分层交叉验证能保证每折中正负样本的比例和原始数据一致,避免因样本分布不均导致的评估偏差。 - 保留独立测试集:交叉验证的目的是调参和评估模型稳定性,最终的泛化能力需要用从未参与过训练/验证的独立测试集来评估,避免数据泄露。
- 最大化训练数据:每折的验证集数据会在其他折中作为训练数据,相当于原来的10%验证集也全部参与了模型训练,提升了数据利用率。
内容的提问来源于stack exchange,提问作者Kirill
相关产品推荐
相关产品推荐

