分层10折交叉验证下,Wandb/Scikit超参数调优逻辑位置疑问
分层10折交叉验证下的超参数调优位置说明
核心原则
超参数调优逻辑绝对不能接触测试集数据,必须完全在训练子集内完成,避免数据泄露影响模型泛化能力。
一、Scikit-learn 工具(GridSearchCV/RandomizedSearchCV)
两种常见场景:
- 直接用调优工具集成分层10折
如果目标是找到最优超参数,同时用分层10折评估调优后的模型性能,直接把分层10折逻辑传给调优工具的cv参数即可,无需手动写外部循环。
示例代码:
from sklearn.model_selection import StratifiedKFold, GridSearchCV from sklearn.svm import SVC # 定义分层10折 skf = StratifiedKFold(n_splits=10, shuffle=True, random_state=42) # 定义模型和超参数空间 model = SVC() param_grid = {'C': [0.1, 1, 10], 'gamma': [1, 0.1, 0.01]} # 把分层10折传给调优工具,内部自动完成调优+交叉验证 grid_search = GridSearchCV(estimator=model, param_grid=param_grid, cv=skf, scoring='accuracy') grid_search.fit(X, y) # 输出最优参数和交叉验证得分 print("最优参数:", grid_search.best_params_) print("交叉验证平均得分:", grid_search.best_score_)
- 手动写分层10折循环,每个折内做超参数调优
如果需要更精细的控制(比如每个折的调优逻辑独立),则把超参数调优放在每个分层折的训练子集内部,外部是分层10折的循环。这种场景下,调优逻辑在循环内部,但仅作用于当前折的训练集。
示例代码:
from sklearn.model_selection import StratifiedKFold, GridSearchCV from sklearn.svm import SVC import numpy as np skf = StratifiedKFold(n_splits=10, shuffle=True, random_state=42) model = SVC() param_grid = {'C': [0.1, 1, 10], 'gamma': [1, 0.1, 0.01]} scores = [] for train_idx, test_idx in skf.split(X, y): X_train, X_test = X[train_idx], X[test_idx] y_train, y_test = y[train_idx], y[test_idx] # 在当前折的训练集内做超参数调优(用内部CV,比如3折) inner_cv = StratifiedKFold(n_splits=3, shuffle=True, random_state=42) grid_search = GridSearchCV(estimator=model, param_grid=param_grid, cv=inner_cv, scoring='accuracy') grid_search.fit(X_train, y_train) # 用最优参数训练模型并评估 best_model = grid_search.best_estimator_ score = best_model.score(X_test, y_test) scores.append(score) print("10折平均得分:", np.mean(scores))
二、Weights & Biases (Wandb)
Wandb的超参数调优(Sweeps)和交叉验证结合时,常用两种方式:
- Sweep 内部集成分层交叉验证
把分层10折的逻辑写在Wandb的训练函数里,Sweep负责遍历超参数,每个超参数组合都会用分层10折评估。这种情况下,交叉验证逻辑在Sweep的训练函数内部,无需外部循环。
示例代码片段:
import wandb from sklearn.model_selection import StratifiedKFold from sklearn.svm import SVC import numpy as np def train(): wandb.init() config = wandb.config skf = StratifiedKFold(n_splits=10, shuffle=True, random_state=42) scores = [] model = SVC(C=config.C, gamma=config.gamma) for train_idx, test_idx in skf.split(X, y): X_train, X_test = X[train_idx], X[test_idx] y_train, y_test = y[train_idx], y[test_idx] model.fit(X_train, y_train) scores.append(model.score(X_test, y_test)) wandb.log({"mean_accuracy": np.mean(scores)}) # 定义Sweep配置 sweep_config = { "method": "grid", "parameters": { "C": {"values": [0.1, 1, 10]}, "gamma": {"values": [1, 0.1, 0.01]} } } sweep_id = wandb.sweep(sweep_config, project="your-project-name") wandb.agent(sweep_id, function=train)
- 外部分层10折循环,每个折内跑Wandb Sweep
这种场景适用于需要每个折的超参数独立调优的情况,此时Sweep的启动逻辑放在分层10折的循环内部,但同样要确保调优仅使用当前折的训练集。不过这种方式较少用,因为会产生大量Sweep任务,一般推荐第一种方式。
总结
- Scikit调优工具:优先用调优工具集成分层10折(逻辑在外部,工具内部处理循环);若需手动控制每个折,就把调优逻辑放在分层折循环的内部训练子集上。
- Wandb Sweep:通常把分层10折逻辑写在Sweep的训练函数内部,Sweep负责遍历超参数,无需额外外部循环。
内容的提问来源于stack exchange,提问作者Ayesha Kiran
相关产品推荐
相关产品推荐

