You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在含OneHot编码的scikit-learn管道中映射XGBRegressor特征重要性与特征名?

解决XGBRegressor特征重要性与管道特征名映射问题

1. 获取管道生成的完整特征列表

训练完管道后,直接调用ColumnTransformer的get_feature_names_out()方法就能得到所有处理后特征的完整列表,包括独热编码后的分类特征和原数值特征,且顺序和模型输入的特征顺序完全一致,不用手动拼接:

# 训练pipeline后执行
full_feature_names = pipeline["encoder"].get_feature_names_out()

这个方法会自动处理:

  • 分类特征:生成带categories__前缀的独热编码特征名(比如categories__cat_feature1_类别A)
  • 数值特征:生成带remainder__前缀的原特征名(如果要去掉前缀,可后续用str.replace('remainder__', '')处理)

2. 映射特征重要性与特征名

XGBoost返回的fscore字典中,键f0、f2等的数字部分对应完整特征列表的索引,只需提取索引匹配特征名即可:

import pandas as pd

# 获取XGB的特征重要性分数
fscore_dict = pipeline['regressor'].get_booster().get_fscore()

# 映射特征名与分数
feature_importance = {}
for key, score in fscore_dict.items():
    # 提取特征索引:去掉键的'f'前缀,转为整数
    feature_idx = int(key[1:])
    # 匹配对应的特征名
    feature_name = full_feature_names[feature_idx]
    feature_importance[feature_name] = score

# 转为DataFrame方便排序和查看
importance_df = pd.DataFrame.from_dict(
    feature_importance, 
    orient='index', 
    columns=['特征重要性(fscore)']
).sort_values(by='特征重要性(fscore)', ascending=False)

print(importance_df)

补充说明

  • 不用手动拼接独热特征和数值特征的原因:get_feature_names_out()严格遵循ColumnTransformer的处理顺序,避免手动拼接时因特征顺序错误导致映射失效的问题
  • 如果需要更简洁的特征名,可对full_feature_names做批量替换,比如去掉所有前缀:
full_feature_names = [name.split('__')[-1] for name in full_feature_names]

内容的提问来源于stack exchange,提问作者Joysn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 17:53:15