如何在含OneHot编码的scikit-learn管道中映射XGBRegressor特征重要性与特征名?
解决XGBRegressor特征重要性与管道特征名映射问题
1. 获取管道生成的完整特征列表
训练完管道后,直接调用ColumnTransformer的get_feature_names_out()方法就能得到所有处理后特征的完整列表,包括独热编码后的分类特征和原数值特征,且顺序和模型输入的特征顺序完全一致,不用手动拼接:
# 训练pipeline后执行 full_feature_names = pipeline["encoder"].get_feature_names_out()
这个方法会自动处理:
- 分类特征:生成带
categories__前缀的独热编码特征名(比如categories__cat_feature1_类别A) - 数值特征:生成带
remainder__前缀的原特征名(如果要去掉前缀,可后续用str.replace('remainder__', '')处理)
2. 映射特征重要性与特征名
XGBoost返回的fscore字典中,键f0、f2等的数字部分对应完整特征列表的索引,只需提取索引匹配特征名即可:
import pandas as pd # 获取XGB的特征重要性分数 fscore_dict = pipeline['regressor'].get_booster().get_fscore() # 映射特征名与分数 feature_importance = {} for key, score in fscore_dict.items(): # 提取特征索引:去掉键的'f'前缀,转为整数 feature_idx = int(key[1:]) # 匹配对应的特征名 feature_name = full_feature_names[feature_idx] feature_importance[feature_name] = score # 转为DataFrame方便排序和查看 importance_df = pd.DataFrame.from_dict( feature_importance, orient='index', columns=['特征重要性(fscore)'] ).sort_values(by='特征重要性(fscore)', ascending=False) print(importance_df)
补充说明
- 不用手动拼接独热特征和数值特征的原因:
get_feature_names_out()严格遵循ColumnTransformer的处理顺序,避免手动拼接时因特征顺序错误导致映射失效的问题 - 如果需要更简洁的特征名,可对
full_feature_names做批量替换,比如去掉所有前缀:
full_feature_names = [name.split('__')[-1] for name in full_feature_names]
内容的提问来源于stack exchange,提问作者Joysn
相关产品推荐
相关产品推荐

