You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

设置随机种子后Python代码跨运行结果仍不可复现的问题

固定随机种子后仍无法复现ShapRFECV结果的问题与解决方案

问题描述

设置完全相同的随机种子、使用静态输入数据的情况下,运行同一Python程序却得到不同结果:Jupyter Notebook同一会话内调用shap_feature_selection函数结果一致,但重启内核后结果改变;命令行运行脚本也存在同样问题。随机性由ShapRFECV引入,现寻求除设置种子外的代码可复现方案。运行环境为仅CPU。

最小复现代码

import os, random
import numpy as np
import pandas as pd
from probatus.feature_elimination import ShapRFECV
from sklearn.ensemble import RandomForestClassifier
from sklearn.datasets import make_classification

global_seed = 1234
os.environ['PYTHONHASHSEED'] = str(global_seed)
np.random.seed(global_seed)
random.seed(global_seed)

feature_names = ['f1', 'f2', 'f3_static', 'f4', 'f5', 'f6', 'f7',
 'f8', 'f9', 'f10', 'f11', 'f12', 'f13', 'f14', 'f15', 'f16', 'f17', 
'f18', 'f19', 'f20']

# Code from tutorial on probatus documentation
X, y = make_classification(n_samples=100, class_sep=0.05, n_informative=6, n_features=20, 
random_state=0, n_redundant=10, n_clusters_per_class=1)
X = pd.DataFrame(X, columns=feature_names)

def shap_feature_selection(X, y, seed: int) -> list[str]:
    
    random_forest = RandomForestClassifier(random_state=seed, n_estimators=70, max_features='log2',
criterion='entropy', class_weight='balanced')
    # Set to run on one thread only
    shap_elimination = ShapRFECV(clf=random_forest, step=0.2, cv=5,
scoring='f1_macro', n_jobs=1, random_state=seed)

    report = shap_elimination.fit_compute(X, y, check_additivity=True, seed=seed)
    # Return the set of features with the best validation accuracy
    return report.iloc[[report['val_metric_mean'].idxmax() - 1]]['features_set'].to_list()[0]

运行结果

# Results from the first run
shap_feature_selection(X, y, 0)

>>> ['f17', 'f15', 'f18', 'f8', 'f12', 'f1', 'f13']

# Running again in same session
shap_feature_selection(X, y, 0)

>>> ['f17', 'f15', 'f18', 'f8', 'f12', 'f1', 'f13']

# Restarting the kernel and running the exact same command
shap_feature_selection(X, y, 0)
>>> ['f8', 'f1', 'f17', 'f6', 'f18', 'f20', 'f12', 'f15', 'f7', 'f13', 'f11']

环境详情

  • Ubuntu 22.04
  • Python 3.9.12
  • Numpy 1.22.0
  • Sklearn 1.1.1

解决方案

1. 调整环境变量与库导入顺序

PYTHONHASHSEED需要在Python解释器初始化阶段生效,必须在导入任何依赖库前设置该环境变量,避免哈希随机化引入的隐性随机性。

2. 显式固定交叉验证拆分器

ShapRFECV默认使用的交叉验证拆分器(如StratifiedKFold)若启用shuffle,未固定种子会导致每次运行的训练/验证拆分不同,需手动创建固定random_state的拆分器并传入。

3. 加固所有随机源控制

确保所有涉及随机操作的组件(随机森林、ShapRFECV、fit_compute方法)都使用相同种子,同时避免其他代码修改随机状态。

修改后的代码示例:

# 先设置环境变量,再导入所有库
import os
global_seed = 1234
os.environ['PYTHONHASHSEED'] = str(global_seed)

import random
import numpy as np
import pandas as pd
from probatus.feature_elimination import ShapRFECV
from sklearn.ensemble import RandomForestClassifier
from sklearn.datasets import make_classification
from sklearn.model_selection import StratifiedKFold

# 固定全局随机种子
np.random.seed(global_seed)
random.seed(global_seed)

feature_names = ['f1', 'f2', 'f3_static', 'f4', 'f5', 'f6', 'f7',
 'f8', 'f9', 'f10', 'f11', 'f12', 'f13', 'f14', 'f15', 'f16', 'f17', 
'f18', 'f19', 'f20']

X, y = make_classification(n_samples=100, class_sep=0.05, n_informative=6, n_features=20, 
random_state=0, n_redundant=10, n_clusters_per_class=1)
X = pd.DataFrame(X, columns=feature_names)

def shap_feature_selection(X, y, seed: int) -> list[str]:
    # 固定随机森林的随机状态
    random_forest = RandomForestClassifier(
        random_state=seed, 
        n_estimators=70, 
        max_features='log2',
        criterion='entropy', 
        class_weight='balanced'
    )
    # 手动创建固定种子的交叉验证拆分器
    cv_splitter = StratifiedKFold(n_splits=5, shuffle=True, random_state=seed)
    
    # 初始化ShapRFECV时传入固定的拆分器和种子
    shap_elimination = ShapRFECV(
        clf=random_forest, 
        step=0.2, 
        cv=cv_splitter,
        scoring='f1_macro', 
        n_jobs=1, 
        random_state=seed
    )

    # fit_compute时也传入种子
    report = shap_elimination.fit_compute(X, y, check_additivity=True, seed=seed)
    # 返回最优特征集
    return report.iloc[[report['val_metric_mean'].idxmax() - 1]]['features_set'].to_list()[0]

额外注意事项

  • 确保运行脚本时没有其他外部进程修改环境变量或随机状态
  • 若使用probatus版本存在隐性随机漏洞,可尝试升级至最新稳定版

内容的提问来源于stack exchange,提问作者Dreana

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 03:30:02