You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

能否为sklearn RFE指定评分指标?如何获取不同特征子集性能

针对类别不平衡场景的RFE特征排序与多子集性能评估方案

方法1:给基础RFE自定义评分逻辑(适配类别不平衡)

基础RFE本身没有scoring参数,但它依赖 estimator 的score方法评估特征子集。我们可以直接包装你的基础模型,重写score方法,替换成适合不平衡数据的指标(比如AUC、F1-score、加权recall等)。

示例代码(以随机森林+AUC为例):

from sklearn.base import clone
from sklearn.ensemble import RandomForestClassifier
from sklearn.feature_selection import RFE
from sklearn.metrics import roc_auc_score
import pandas as pd

# 包装模型,重写score方法为AUC
class ScoredEstimator(RandomForestClassifier):
    def score(self, X, y, sample_weight=None):
        y_pred_proba = self.predict_proba(X)[:, 1]
        return roc_auc_score(y, y_pred_proba, sample_weight=sample_weight)

# 初始化自定义模型
custom_estimator = ScoredEstimator(n_estimators=100, random_state=42)
feature_names = X.columns.tolist()
results = []

# 循环生成Top1、Top2、Top3等特征子集
for k in [1, 2, 3]:
    rfe = RFE(estimator=custom_estimator, n_features_to_select=k, step=1)
    rfe.fit(X, y)
    # 获取选中的特征
    selected_features = [feature_names[i] for i, mask in enumerate(rfe.support_) if mask]
    # 用自定义评分获取性能
    performance = round(rfe.score(X, y), 2)
    results.append({
        "Number of features": k,
        "Features": ", ".join(selected_features),
        "Performance": performance
    })

# 输出结果
print(pd.DataFrame(results))

这种方式让RFE在迭代剔除特征时,全程用你指定的不平衡友好指标评估,同时能快速生成不同k值的特征子集。

方法2:改造RFECV获取全范围特征排名

RFECV的ranking_仅将最优特征子集标记为1,但我们可以手动模拟RFE的迭代逻辑,结合RFECV支持的scoring参数,生成完整的特征重要性排名:

from sklearn.metrics import roc_auc_score
from sklearn.inspection import permutation_importance
from sklearn.ensemble import RandomForestClassifier
import numpy as np
import pandas as pd

def custom_rfe_with_scoring(X, y, estimator, scoring_func, max_k):
    feature_names = X.columns.tolist()
    current_features = feature_names.copy()
    performance_dict = {}
    
    for k in range(max_k, 0, -1):
        # 训练模型并评估当前特征子集性能
        est = clone(estimator)
        est.fit(X[current_features], y)
        # 根据评分函数类型计算性能
        if scoring_func.__name__ == "roc_auc_score":
            y_pred_proba = est.predict_proba(X[current_features])[:, 1]
            perf = scoring_func(y, y_pred_proba)
        else:
            y_pred = est.predict(X[current_features])
            perf = scoring_func(y, y_pred)
        performance_dict[k] = (current_features.copy(), round(perf, 2))
        
        if k == 1:
            break
        
        # 用排列重要性评估单个特征贡献(比模型自带重要性更适配不平衡数据)
        perm_result = permutation_importance(est, X[current_features], y, scoring=scoring_func, n_repeats=5, random_state=42)
        # 找到重要性最低的特征并剔除
        least_important_idx = np.argmin(perm_result.importances_mean)
        current_features.pop(least_important_idx)
    
    # 整理为Top k格式的结果
    results = []
    for k in range(1, max_k+1):
        feats, perf = performance_dict[k]
        results.append({
            "Number of features": k,
            "Features": ", ".join(feats),
            "Performance": perf
        })
    return results

# 使用示例
estimator = RandomForestClassifier(n_estimators=100, random_state=42)
results = custom_rfe_with_scoring(X, y, estimator, roc_auc_score, max_k=3)
print(pd.DataFrame(results))

这里用排列重要性替代模型自带的特征重要性,避免不平衡数据带来的偏差,同时全程用指定评分指标评估。

关于SFS的替代建议

SFS对500特征的效率极低,完全没必要硬用。你可以先通过SelectFromModel筛选掉低重要性特征(比如保留前200个),再用上述RFE方法处理,既能提速,又不会大幅损失性能。

内容的提问来源于stack exchange,提问作者Charlie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 17:25:21