You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何替换train_test_split验证采样,转为全交叉验证方案?

改用全交叉验证的实现方案

原来的固定拆分方式会占用10%的数据作为独立验证集,无法参与模型训练。改用分层k折交叉验证后,所有非测试集的数据都会轮流作为训练集和验证集,最大化训练数据量,同时保持验证评估的可靠性。

代码实现

from sklearn.model_selection import StratifiedKFold, train_test_split

# 提取特征和标签
X = data['text_with_tokeniz_lemmatiz']
y = data['toxic']

# 先拆分出10%的独立测试集(仅用于最终评估)
X_train_val, X_test, y_train_val, y_test = train_test_split(
    X, y, test_size=0.1, shuffle=True, random_state=12345, stratify=y
)

# 初始化9折分层交叉验证(保证每折类别分布与整体一致)
skf = StratifiedKFold(n_splits=9, shuffle=True, random_state=12345)

# 遍历每折进行训练与验证
for fold_idx, (train_indices, val_indices) in enumerate(skf.split(X_train_val, y_train_val)):
    # 划分当前折的训练集和验证集
    X_train, X_val = X_train_val.iloc[train_indices], X_train_val.iloc[val_indices]
    y_train, y_val = y_train_val.iloc[train_indices], y_train_val.iloc[val_indices]
    
    # 替换为你的模型训练代码
    # model = 你的模型实例
    # model.fit(X_train, y_train)
    # val_performance = model.evaluate(X_val, y_val)
    print(f"第{fold_idx+1}折,验证集性能:{val_performance}")

# 用全部训练验证数据训练最终模型,再评估测试集
final_model.fit(X_train_val, y_train_val)
test_performance = final_model.evaluate(X_test, y_test)
print(f"最终测试集性能:{test_performance}")

关键说明

  • 选择StratifiedKFold而非普通KFold:因为你的任务是二分类(toxic标签),分层交叉验证能保证每折中正负样本的比例和原始数据一致,避免因样本分布不均导致的评估偏差。
  • 保留独立测试集:交叉验证的目的是调参和评估模型稳定性,最终的泛化能力需要用从未参与过训练/验证的独立测试集来评估,避免数据泄露。
  • 最大化训练数据:每折的验证集数据会在其他折中作为训练数据,相当于原来的10%验证集也全部参与了模型训练,提升了数据利用率。

内容的提问来源于stack exchange,提问作者Kirill

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 10:05:30