设置随机种子后Python代码跨运行结果仍不可复现的问题
固定随机种子后仍无法复现ShapRFECV结果的问题与解决方案
问题描述
设置完全相同的随机种子、使用静态输入数据的情况下,运行同一Python程序却得到不同结果:Jupyter Notebook同一会话内调用shap_feature_selection函数结果一致,但重启内核后结果改变;命令行运行脚本也存在同样问题。随机性由ShapRFECV引入,现寻求除设置种子外的代码可复现方案。运行环境为仅CPU。
最小复现代码
import os, random import numpy as np import pandas as pd from probatus.feature_elimination import ShapRFECV from sklearn.ensemble import RandomForestClassifier from sklearn.datasets import make_classification global_seed = 1234 os.environ['PYTHONHASHSEED'] = str(global_seed) np.random.seed(global_seed) random.seed(global_seed) feature_names = ['f1', 'f2', 'f3_static', 'f4', 'f5', 'f6', 'f7', 'f8', 'f9', 'f10', 'f11', 'f12', 'f13', 'f14', 'f15', 'f16', 'f17', 'f18', 'f19', 'f20'] # Code from tutorial on probatus documentation X, y = make_classification(n_samples=100, class_sep=0.05, n_informative=6, n_features=20, random_state=0, n_redundant=10, n_clusters_per_class=1) X = pd.DataFrame(X, columns=feature_names) def shap_feature_selection(X, y, seed: int) -> list[str]: random_forest = RandomForestClassifier(random_state=seed, n_estimators=70, max_features='log2', criterion='entropy', class_weight='balanced') # Set to run on one thread only shap_elimination = ShapRFECV(clf=random_forest, step=0.2, cv=5, scoring='f1_macro', n_jobs=1, random_state=seed) report = shap_elimination.fit_compute(X, y, check_additivity=True, seed=seed) # Return the set of features with the best validation accuracy return report.iloc[[report['val_metric_mean'].idxmax() - 1]]['features_set'].to_list()[0]
运行结果
# Results from the first run shap_feature_selection(X, y, 0) >>> ['f17', 'f15', 'f18', 'f8', 'f12', 'f1', 'f13'] # Running again in same session shap_feature_selection(X, y, 0) >>> ['f17', 'f15', 'f18', 'f8', 'f12', 'f1', 'f13'] # Restarting the kernel and running the exact same command shap_feature_selection(X, y, 0) >>> ['f8', 'f1', 'f17', 'f6', 'f18', 'f20', 'f12', 'f15', 'f7', 'f13', 'f11']
环境详情
- Ubuntu 22.04
- Python 3.9.12
- Numpy 1.22.0
- Sklearn 1.1.1
解决方案
1. 调整环境变量与库导入顺序
PYTHONHASHSEED需要在Python解释器初始化阶段生效,必须在导入任何依赖库前设置该环境变量,避免哈希随机化引入的隐性随机性。
2. 显式固定交叉验证拆分器
ShapRFECV默认使用的交叉验证拆分器(如StratifiedKFold)若启用shuffle,未固定种子会导致每次运行的训练/验证拆分不同,需手动创建固定random_state的拆分器并传入。
3. 加固所有随机源控制
确保所有涉及随机操作的组件(随机森林、ShapRFECV、fit_compute方法)都使用相同种子,同时避免其他代码修改随机状态。
修改后的代码示例:
# 先设置环境变量,再导入所有库 import os global_seed = 1234 os.environ['PYTHONHASHSEED'] = str(global_seed) import random import numpy as np import pandas as pd from probatus.feature_elimination import ShapRFECV from sklearn.ensemble import RandomForestClassifier from sklearn.datasets import make_classification from sklearn.model_selection import StratifiedKFold # 固定全局随机种子 np.random.seed(global_seed) random.seed(global_seed) feature_names = ['f1', 'f2', 'f3_static', 'f4', 'f5', 'f6', 'f7', 'f8', 'f9', 'f10', 'f11', 'f12', 'f13', 'f14', 'f15', 'f16', 'f17', 'f18', 'f19', 'f20'] X, y = make_classification(n_samples=100, class_sep=0.05, n_informative=6, n_features=20, random_state=0, n_redundant=10, n_clusters_per_class=1) X = pd.DataFrame(X, columns=feature_names) def shap_feature_selection(X, y, seed: int) -> list[str]: # 固定随机森林的随机状态 random_forest = RandomForestClassifier( random_state=seed, n_estimators=70, max_features='log2', criterion='entropy', class_weight='balanced' ) # 手动创建固定种子的交叉验证拆分器 cv_splitter = StratifiedKFold(n_splits=5, shuffle=True, random_state=seed) # 初始化ShapRFECV时传入固定的拆分器和种子 shap_elimination = ShapRFECV( clf=random_forest, step=0.2, cv=cv_splitter, scoring='f1_macro', n_jobs=1, random_state=seed ) # fit_compute时也传入种子 report = shap_elimination.fit_compute(X, y, check_additivity=True, seed=seed) # 返回最优特征集 return report.iloc[[report['val_metric_mean'].idxmax() - 1]]['features_set'].to_list()[0]
额外注意事项
- 确保运行脚本时没有其他外部进程修改环境变量或随机状态
- 若使用probatus版本存在隐性随机漏洞,可尝试升级至最新稳定版
内容的提问来源于stack exchange,提问作者Dreana
相关产品推荐
相关产品推荐

