You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Sklearn Pipeline中提取带特征名的模型特征重要性

在含预处理的Sklearn Pipeline中提取带特征名的特征重要性(原生方法vs ELI5)

需求说明

需要从包含MinMaxScaler预处理步骤的Sklearn Pipeline中,为Logistic Regression、Gradient Boosting(GBM)、XGBoost三个分类器提取带原始特征名称的特征重要性,并对比Sklearn原生方法与ELI5工具的结果差异。

基础训练代码

以下是完整的模型训练与评估代码(修正了原代码中的语法细节):

import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split, cross_validate, StratifiedKFold
from sklearn.preprocessing import MinMaxScaler
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import GradientBoostingClassifier
from sklearn.metrics import classification_report, confusion_matrix
from xgboost import XGBClassifier
import matplotlib.pyplot as plt
import seaborn as sns

# 定义特征与目标变量
X = df3.drop(['Prediction_SAP_Burst','Unnamed: 0'], axis=1)
y = df3['Prediction_SAP_Burst']

# 划分训练测试集
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=.30, random_state=1, stratify=y)
print(X_train.shape)
print(X_test.shape)
print(y_train.shape)
print(y_test.shape)

# 定义数值特征与预处理管道(保留特征顺序,避免set打乱)
numerical_features = X.columns.tolist()
numerical_transformer = Pipeline(steps=[('scaler', MinMaxScaler())])

# 构建预处理处理器
preprocessor = ColumnTransformer(
    transformers=[('num', numerical_transformer, numerical_features)]
)

# 创建各模型的Pipeline
lr_clf = Pipeline(steps=[
    ('preprocessor', preprocessor),
    ('lr', LogisticRegression(class_weight={0:0.52,1:16.14}, random_state=1))
])

xgb_clf = Pipeline(steps=[
    ('preprocessor', preprocessor),
    ('classifier', XGBClassifier(scale_pos_weight=8, random_state=1))
])

gbc_clf = Pipeline(steps=[
    ('preprocessor', preprocessor),
    ('classifier', GradientBoostingClassifier(
        random_state=1, learning_rate=0.2, max_features=0.5,
        n_estimators=250, subsample=0.8
    ))
])

# 训练模型
lr_clf.fit(X_train, y_train)
xgb_clf.fit(X_train, y_train)
gbc_clf.fit(X_train, y_train)

# 模型评估函数
def results(name: str, model) -> None:
    preds = model.predict(X_test)
    model_cv = cross_validate(model, X_train, y_train, cv=StratifiedKFold(n_splits=5), n_jobs=-1, scoring='f1')
    print(f"Kfold F1分数: {model_cv['test_score']}")
    print(f"Kfold平均F1分数: {model_cv['test_score'].mean():.3f} +/- {model_cv['test_score'].std():.3f}")
    print(f"{name} 测试集准确率: {model.score(X_test, y_test):.3f}")
    print(classification_report(y_test, preds))
    
    # 绘制混淆矩阵
    labels = ['Good', 'Bad']
    conf_matrix = confusion_matrix(y_test, preds, normalize='true')
    plt.rc('font', family='normal', size=14)
    plt.figure(figsize=(10,6))
    sns.heatmap(conf_matrix, xticklabels=labels, yticklabels=labels, annot=True, cmap='Blues')
    plt.title(f"{name} 混淆矩阵")
    plt.ylabel('真实类别')
    plt.xlabel('预测类别')
    plt.show()

# 执行评估
results("逻辑回归", lr_clf)
results("XGBoost", xgb_clf)
results("梯度提升树(GBM)", gbc_clf)

一、Sklearn原生方法提取特征重要性

Sklearn原生支持通过模型属性获取特征重要性/系数,结合ColumnTransformer.get_feature_names_out()可精准关联原始特征名称:

通用提取函数

def get_feature_importance_with_names(pipeline, feature_names):
    # 获取预处理后的特征名(此处与原始特征名一致)
    processed_feature_names = pipeline.named_steps['preprocessor'].get_feature_names_out(feature_names)
    
    # 提取Pipeline中的模型
    model = pipeline.named_steps[next(key for key in pipeline.named_steps if key != 'preprocessor')]
    
    # 根据模型类型获取重要性数据
    if isinstance(model, LogisticRegression):
        # 二分类问题取第一类的系数
        importance = model.coef_[0]
    elif isinstance(model, (GradientBoostingClassifier, XGBClassifier)):
        # 树模型取特征重要性
        importance = model.feature_importances_
    else:
        raise ValueError("不支持当前模型类型")
    
    # 组合成DataFrame并按重要性绝对值排序
    feature_importance_df = pd.DataFrame({
        '特征名称': processed_feature_names,
        '重要性': importance
    }).sort_values(by='重要性', key=abs, ascending=False)
    
    return feature_importance_df

# 提取各模型的特征重要性
lr_importance = get_feature_importance_with_names(lr_clf, numerical_features)
xgb_importance = get_feature_importance_with_names(xgb_clf, numerical_features)
gbc_importance = get_feature_importance_with_names(gbc_clf, numerical_features)

# 打印前10个重要特征
print("逻辑回归特征重要性:")
print(lr_importance.head(10))
print("\nXGBoost特征重要性:")
print(xgb_importance.head(10))
print("\nGBM特征重要性:")
print(gbc_importance.head(10))

注意事项

  • get_feature_importance_with_names()是Sklearn 1.0+版本新增方法,能准确返回预处理后的特征名,避免手动匹配错误;
  • 逻辑回归的coef_基于缩放后的数据,系数大小反映缩放后特征对预测的影响;
  • 树模型的feature_importances_基于特征分裂增益,与预处理无关(树模型对特征缩放不敏感)。

二、ELI5方法提取特征重要性

ELI5可直接处理Sklearn Pipeline,自动跳过预处理步骤,一键返回带原始特征名的重要性:

安装与使用

先安装ELI5:

pip install eli5

然后编写提取代码:

import eli5

# 逻辑回归:显示特征权重
print("ELI5 逻辑回归特征权重:")
display(eli5.show_weights(lr_clf, feature_names=numerical_features, top=10))

# XGBoost:显示特征重要性
print("\nELI5 XGBoost特征重要性:")
display(eli5.show_weights(xgb_clf, feature_names=numerical_features, top=10))

# GBM:显示特征重要性
print("\nELI5 GBM特征重要性:")
display(eli5.show_weights(gbc_clf, feature_names=numerical_features, top=10))

注意事项

  • ELI5对Logistic Regression默认显示标准化后的权重,更易对比特征间的相对影响;
  • 对树模型,ELI5的特征重要性数值与原生feature_importances_完全一致,展示形式更直观;
  • 支持在Jupyter Notebook中直接可视化,无需手动转换格式。

三、结果对比

维度Sklearn原生方法ELI5方法
操作复杂度需手动关联特征名与重要性,代码量略多一键生成结果,无需额外处理
逻辑回归结果返回缩放后数据的系数,需自行解读趋势返回标准化权重,特征间影响对比更清晰
树模型结果与ELI5数值完全一致与原生数值一致,展示形式更友好
可视化支持需手动绘制(如barplot)内置可视化,直接在Notebook展示

总结

  • Sklearn原生方法适合需要自定义处理流程的场景,完全满足需求;
  • ELI5更便捷,快速分析与可视化时可读性更强;
  • 两种方法对树模型的特征重要性结果完全一致,逻辑回归结果趋势一致,数值差异源于标准化/缩放处理。

内容的提问来源于stack exchange,提问作者MOT

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 13:50:33