You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

能否在Learning API中获取XGBoost与LightGBM的排列重要性?

获取XGBoost/LightGBM Learning API模型的排列重要性

排列重要性的核心逻辑是:打乱单个特征的取值后,观察模型性能的下降幅度——下降越多,说明该特征对模型的预测能力越重要。针对Learning API训练的模型(即xgboost.train/lightgbm.train得到的Booster对象),可以通过两种方式实现:

方法一:手动实现排列重要性

XGBoost Learning API 示例

假设你已经用Learning API训练好XGBoost模型,以下是二分类任务下的手动实现代码:

import numpy as np
import xgboost as xgb
from sklearn.metrics import log_loss

# 1. 准备数据并训练Learning API模型
X_train, y_train = ... # 替换为你的训练数据
X_val, y_val = ... # 替换为你的验证数据(用于计算重要性)
dtrain = xgb.DMatrix(X_train, label=y_train)
dval = xgb.DMatrix(X_val, label=y_val)

params = {"objective": "binary:logistic", "eval_metric": "logloss"}
booster = xgb.train(params, dtrain, num_boost_round=100, evals=[(dval, "val")])

# 2. 定义性能评估函数(这里用二分类的logloss)
def evaluate_model(booster, X, y):
    dmat = xgb.DMatrix(X)
    y_pred = booster.predict(dmat)
    return log_loss(y, y_pred)

# 3. 计算基准性能
baseline_score = evaluate_model(booster, X_val, y_val)
feature_names = X_val.columns.tolist()
perm_importances = []

# 4. 遍历每个特征计算排列重要性
for feature in feature_names:
    # 复制数据并打乱当前特征
    X_perm = X_val.copy()
    X_perm[feature] = np.random.permutation(X_perm[feature])
    # 计算打乱后的性能
    perm_score = evaluate_model(booster, X_perm, y_val)
    # 重要性为性能下降值(越大越重要)
    importance = baseline_score - perm_score
    perm_importances.append(importance)

# 整理并输出结果
perm_importance_dict = dict(zip(feature_names, perm_importances))
sorted_importances = sorted(perm_importance_dict.items(), key=lambda x: x[1], reverse=True)
print("XGBoost Learning API 排列重要性:")
for feat, imp in sorted_importances:
    print(f"{feat}: {imp:.4f}")

LightGBM Learning API 示例

逻辑与XGBoost一致,针对LightGBM的Booster对象实现:

import numpy as np
import lightgbm as lgb
from sklearn.metrics import log_loss

# 1. 准备数据并训练Learning API模型
X_train, y_train = ... # 替换为你的训练数据
X_val, y_val = ... # 替换为你的验证数据
lgb_train = lgb.Dataset(X_train, label=y_train)
lgb_val = lgb.Dataset(X_val, label=y_val, reference=lgb_train)

params = {"objective": "binary", "metric": "binary_logloss"}
booster = lgb.train(params, lgb_train, num_boost_round=100, valid_sets=[lgb_val])

# 2. 定义性能评估函数
def evaluate_model(booster, X, y):
    y_pred = booster.predict(X, num_iteration=booster.best_iteration)
    return log_loss(y, y_pred)

# 3. 计算基准性能
baseline_score = evaluate_model(booster, X_val, y_val)
feature_names = X_val.columns.tolist()
perm_importances = []

# 4. 遍历特征计算重要性
for feature in feature_names:
    X_perm = X_val.copy()
    X_perm[feature] = np.random.permutation(X_perm[feature])
    perm_score = evaluate_model(booster, X_perm, y_val)
    importance = baseline_score - perm_score
    perm_importances.append(importance)

# 整理并输出结果
perm_importance_dict = dict(zip(feature_names, perm_importances))
sorted_importances = sorted(perm_importance_dict.items(), key=lambda x: x[1], reverse=True)
print("LightGBM Learning API 排列重要性:")
for feat, imp in sorted_importances:
    print(f"{feat}: {imp:.4f}")

方法二:包装模型适配sklearn的permutation_importance

sklearn的sklearn.inspection.permutation_importance只兼容符合sklearn接口规范的模型(即有predict/predict_proba方法)。我们可以写一个简单的包装类,把Learning API的Booster对象转换成sklearn兼容的estimator,直接复用sklearn的工具:

XGBoost Booster 包装示例

import xgboost as xgb
from sklearn.inspection import permutation_importance
import numpy as np

class XGBoosterWrapper:
    def __init__(self, booster, y_train):
        self.booster = booster
        self.classes_ = np.unique(y_train)
    
    def predict_proba(self, X):
        dmat = xgb.DMatrix(X)
        pred = self.booster.predict(dmat)
        # 二分类时调整输出形状适配sklearn
        return pred.reshape(-1, 1) if len(self.classes_) == 2 else pred
    
    def predict(self, X):
        proba = self.predict_proba(X)
        return np.argmax(proba, axis=1) if len(proba.shape) > 1 else (proba > 0.5).astype(int)

# 训练好的XGBoost Booster对象
booster = ... 

# 包装模型
sklearn_compatible_model = XGBoosterWrapper(booster, y_train)

# 调用sklearn的permutation_importance
result = permutation_importance(
    sklearn_compatible_model, X_val, y_val,
    scoring="neg_log_loss", n_repeats=10, random_state=42
)

# 整理并输出结果
sorted_idx = result.importances_mean.argsort()[::-1]
print("XGBoost 排列重要性(sklearn工具):")
for idx in sorted_idx:
    print(f"{feature_names[idx]}: {result.importances_mean[idx]:.4f} (±{result.importances_std[idx]:.4f})")

LightGBM Booster 包装示例

import lightgbm as lgb
from sklearn.inspection import permutation_importance
import numpy as np

class LGBMBoosterWrapper:
    def __init__(self, booster, y_train):
        self.booster = booster
        self.classes_ = np.unique(y_train)
    
    def predict_proba(self, X):
        pred = self.booster.predict(X, num_iteration=self.booster.best_iteration)
        # 二分类时调整输出形状适配sklearn
        return pred.reshape(-1, 1) if len(self.classes_) == 2 else pred
    
    def predict(self, X):
        proba = self.predict_proba(X)
        return np.argmax(proba, axis=1) if len(proba.shape) > 1 else (proba > 0.5).astype(int)

# 训练好的LightGBM Booster对象
booster = ... 

# 包装模型
sklearn_compatible_model = LGBMBoosterWrapper(booster, y_train)

# 调用sklearn的permutation_importance
result = permutation_importance(
    sklearn_compatible_model, X_val, y_val,
    scoring="neg_log_loss", n_repeats=10, random_state=42
)

# 整理并输出结果
sorted_idx = result.importances_mean.argsort()[::-1]
print("LightGBM 排列重要性(sklearn工具):")
for idx in sorted_idx:
    print(f"{feature_names[idx]}: {result.importances_mean[idx]:.4f} (±{result.importances_std[idx]:.4f})")

注意:

  • 上述代码中的scoring参数需要根据任务类型调整(比如回归任务用neg_mean_squared_error)。
  • 设置n_repeats多次重复打乱可以让重要性结果更稳定。

内容的提问来源于stack exchange,提问作者royjp

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 22:39:17