使用Optuna调参时同random_state却结果不同的原因排查
为什么设置random_state=42后Optuna运行结果仍有随机性?
你代码里虽然给RandomForestClassifier设置了random_state=42,但每次运行的最优参数仍不一致,核心是还有两处随机性未被固定:
1. Optuna的超参数采样过程未固定种子
Optuna默认使用的TPESampler(树形结构Parzen估计器采样器)本身带有随机性,它会根据之前trial的结果动态采样下一组超参数。如果不固定Optuna的随机种子,每次运行时trial.suggest_int和trial.suggest_float生成的参数组合序列就会不同——你只跑3次trial,不同的参数组合自然可能产生不同的“最优”结果。
解决办法:
- 全局设置Optuna种子:在代码开头添加
optuna.set_seed(42) - 或者创建study时指定采样器的种子:
study = optuna.create_study( direction="maximize", sampler=optuna.samplers.TPESampler(seed=42) )
2. 交叉验证的数据集拆分(若开启shuffle)
如果你的交叉验证开启了shuffle=True(实际调优场景中常用),sklearn的KFold拆分过程会引入随机性。即使模型本身固定了种子,不同的训练/测试拆分也会导致交叉验证得分有差异,进而影响Optuna对参数优劣的判断。
解决办法:显式创建KFold对象并固定随机种子:
from sklearn.model_selection import KFold cv = KFold(n_splits=3, shuffle=True, random_state=42) score = cross_val_score(rf_model, x, y, cv=cv).mean()
验证修改后的代码
固定Optuna种子后,每次运行的参数采样序列会完全一致,3次trial的参数组合相同,最终最优参数也会固定:
import optuna import sklearn from sklearn.ensemble import RandomForestClassifier from sklearn.model_selection import cross_val_score # 固定Optuna全局种子 optuna.set_seed(42) def objective(trial): digits = sklearn.datasets.load_digits() x, y = digits.data, digits.target max_depth = trial.suggest_int("rf_max_depth", 2, 64, log=True) max_samples = trial.suggest_float("rf_max_samples", 0.2, 1) rf_model = RandomForestClassifier( max_depth = max_depth, max_samples = max_samples, n_estimators = 50, random_state = 42 ) score = cross_val_score(rf_model, x, y, cv=3).mean() return score study = optuna.create_study(direction = "maximize") study.optimize(objective, n_trials = 3) trial = study.best_trial print("Best Score: ", trial.value) print("Best Params: ") for key, value in trial.params.items(): print(" {}: {}".format(key, value))
内容的提问来源于stack exchange,提问作者shsh
相关产品推荐
相关产品推荐

