如何从Sklearn Pipeline中提取带特征名的模型特征重要性
在含预处理的Sklearn Pipeline中提取带特征名的特征重要性(原生方法vs ELI5)
需求说明
需要从包含MinMaxScaler预处理步骤的Sklearn Pipeline中,为Logistic Regression、Gradient Boosting(GBM)、XGBoost三个分类器提取带原始特征名称的特征重要性,并对比Sklearn原生方法与ELI5工具的结果差异。
基础训练代码
以下是完整的模型训练与评估代码(修正了原代码中的语法细节):
import pandas as pd import numpy as np from sklearn.model_selection import train_test_split, cross_validate, StratifiedKFold from sklearn.preprocessing import MinMaxScaler from sklearn.compose import ColumnTransformer from sklearn.pipeline import Pipeline from sklearn.linear_model import LogisticRegression from sklearn.ensemble import GradientBoostingClassifier from sklearn.metrics import classification_report, confusion_matrix from xgboost import XGBClassifier import matplotlib.pyplot as plt import seaborn as sns # 定义特征与目标变量 X = df3.drop(['Prediction_SAP_Burst','Unnamed: 0'], axis=1) y = df3['Prediction_SAP_Burst'] # 划分训练测试集 X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=.30, random_state=1, stratify=y) print(X_train.shape) print(X_test.shape) print(y_train.shape) print(y_test.shape) # 定义数值特征与预处理管道(保留特征顺序,避免set打乱) numerical_features = X.columns.tolist() numerical_transformer = Pipeline(steps=[('scaler', MinMaxScaler())]) # 构建预处理处理器 preprocessor = ColumnTransformer( transformers=[('num', numerical_transformer, numerical_features)] ) # 创建各模型的Pipeline lr_clf = Pipeline(steps=[ ('preprocessor', preprocessor), ('lr', LogisticRegression(class_weight={0:0.52,1:16.14}, random_state=1)) ]) xgb_clf = Pipeline(steps=[ ('preprocessor', preprocessor), ('classifier', XGBClassifier(scale_pos_weight=8, random_state=1)) ]) gbc_clf = Pipeline(steps=[ ('preprocessor', preprocessor), ('classifier', GradientBoostingClassifier( random_state=1, learning_rate=0.2, max_features=0.5, n_estimators=250, subsample=0.8 )) ]) # 训练模型 lr_clf.fit(X_train, y_train) xgb_clf.fit(X_train, y_train) gbc_clf.fit(X_train, y_train) # 模型评估函数 def results(name: str, model) -> None: preds = model.predict(X_test) model_cv = cross_validate(model, X_train, y_train, cv=StratifiedKFold(n_splits=5), n_jobs=-1, scoring='f1') print(f"Kfold F1分数: {model_cv['test_score']}") print(f"Kfold平均F1分数: {model_cv['test_score'].mean():.3f} +/- {model_cv['test_score'].std():.3f}") print(f"{name} 测试集准确率: {model.score(X_test, y_test):.3f}") print(classification_report(y_test, preds)) # 绘制混淆矩阵 labels = ['Good', 'Bad'] conf_matrix = confusion_matrix(y_test, preds, normalize='true') plt.rc('font', family='normal', size=14) plt.figure(figsize=(10,6)) sns.heatmap(conf_matrix, xticklabels=labels, yticklabels=labels, annot=True, cmap='Blues') plt.title(f"{name} 混淆矩阵") plt.ylabel('真实类别') plt.xlabel('预测类别') plt.show() # 执行评估 results("逻辑回归", lr_clf) results("XGBoost", xgb_clf) results("梯度提升树(GBM)", gbc_clf)
一、Sklearn原生方法提取特征重要性
Sklearn原生支持通过模型属性获取特征重要性/系数,结合ColumnTransformer.get_feature_names_out()可精准关联原始特征名称:
通用提取函数
def get_feature_importance_with_names(pipeline, feature_names): # 获取预处理后的特征名(此处与原始特征名一致) processed_feature_names = pipeline.named_steps['preprocessor'].get_feature_names_out(feature_names) # 提取Pipeline中的模型 model = pipeline.named_steps[next(key for key in pipeline.named_steps if key != 'preprocessor')] # 根据模型类型获取重要性数据 if isinstance(model, LogisticRegression): # 二分类问题取第一类的系数 importance = model.coef_[0] elif isinstance(model, (GradientBoostingClassifier, XGBClassifier)): # 树模型取特征重要性 importance = model.feature_importances_ else: raise ValueError("不支持当前模型类型") # 组合成DataFrame并按重要性绝对值排序 feature_importance_df = pd.DataFrame({ '特征名称': processed_feature_names, '重要性': importance }).sort_values(by='重要性', key=abs, ascending=False) return feature_importance_df # 提取各模型的特征重要性 lr_importance = get_feature_importance_with_names(lr_clf, numerical_features) xgb_importance = get_feature_importance_with_names(xgb_clf, numerical_features) gbc_importance = get_feature_importance_with_names(gbc_clf, numerical_features) # 打印前10个重要特征 print("逻辑回归特征重要性:") print(lr_importance.head(10)) print("\nXGBoost特征重要性:") print(xgb_importance.head(10)) print("\nGBM特征重要性:") print(gbc_importance.head(10))
注意事项
get_feature_importance_with_names()是Sklearn 1.0+版本新增方法,能准确返回预处理后的特征名,避免手动匹配错误;- 逻辑回归的
coef_基于缩放后的数据,系数大小反映缩放后特征对预测的影响; - 树模型的
feature_importances_基于特征分裂增益,与预处理无关(树模型对特征缩放不敏感)。
二、ELI5方法提取特征重要性
ELI5可直接处理Sklearn Pipeline,自动跳过预处理步骤,一键返回带原始特征名的重要性:
安装与使用
先安装ELI5:
pip install eli5
然后编写提取代码:
import eli5 # 逻辑回归:显示特征权重 print("ELI5 逻辑回归特征权重:") display(eli5.show_weights(lr_clf, feature_names=numerical_features, top=10)) # XGBoost:显示特征重要性 print("\nELI5 XGBoost特征重要性:") display(eli5.show_weights(xgb_clf, feature_names=numerical_features, top=10)) # GBM:显示特征重要性 print("\nELI5 GBM特征重要性:") display(eli5.show_weights(gbc_clf, feature_names=numerical_features, top=10))
注意事项
- ELI5对Logistic Regression默认显示标准化后的权重,更易对比特征间的相对影响;
- 对树模型,ELI5的特征重要性数值与原生
feature_importances_完全一致,展示形式更直观; - 支持在Jupyter Notebook中直接可视化,无需手动转换格式。
三、结果对比
| 维度 | Sklearn原生方法 | ELI5方法 |
|---|---|---|
| 操作复杂度 | 需手动关联特征名与重要性,代码量略多 | 一键生成结果,无需额外处理 |
| 逻辑回归结果 | 返回缩放后数据的系数,需自行解读趋势 | 返回标准化权重,特征间影响对比更清晰 |
| 树模型结果 | 与ELI5数值完全一致 | 与原生数值一致,展示形式更友好 |
| 可视化支持 | 需手动绘制(如barplot) | 内置可视化,直接在Notebook展示 |
总结
- Sklearn原生方法适合需要自定义处理流程的场景,完全满足需求;
- ELI5更便捷,快速分析与可视化时可读性更强;
- 两种方法对树模型的特征重要性结果完全一致,逻辑回归结果趋势一致,数值差异源于标准化/缩放处理。
内容的提问来源于stack exchange,提问作者MOT
相关产品推荐
相关产品推荐

