XGBoost plot_importance()特征重要性数值格式化技术问询
解决XGBoost特征重要性图数值显示拥挤的问题
我之前也碰到过这个坑——xgb.plot_importance()默认显示的重要性值小数位数太多,挤在一起看着特别难受。其实不用找什么隐藏参数,咱们直接利用matplotlib的轴对象就能手动调整数值格式,给你两种靠谱的解决办法:
方法一:快速修改现有绘图的数值显示
plot_importance()返回的ax对象里,那些显示在条形右侧的数值都是独立的文本元素,咱们只需要遍历这些元素,把数值格式化成整数(或者你想要的小数位数)就行:
xg_reg = xgb.XGBClassifier( objective = 'binary:logistic', colsample_bytree = 0.4, learning_rate = 0.01, max_depth = 15, alpha = 0.1, n_estimators = 5, subsample = 0.5, scale_pos_weight = 4 ) xg_reg.fit(X_train, y_train) preds = xg_reg.predict(X_test) # 先画出默认的特征重要性图 ax = xgb.plot_importance(xg_reg, max_num_features=3, importance_type='gain', show_values=True) fig = ax.figure fig.set_size_inches(10, 3) # 遍历所有数值文本,格式化为整数 for text in ax.texts: current_val = float(text.get_text()) # 改成整数显示 text.set_text(f'{int(current_val)}') # 如果想保留1位小数,就写成 f'{current_val:.1f}'
运行这段代码后,原来的"25.66521"就会变成"25",瞬间清爽很多~
方法二:自定义绘图(更灵活可控)
如果你想完全掌控绘图细节(比如调整数值位置、字体大小、条形颜色),也可以自己提取特征重要性数据,用matplotlib手动绘制条形图:
import matplotlib.pyplot as plt # 从模型中提取特征重要性数据(gain类型) importance_dict = xg_reg.get_booster().get_score(importance_type='gain') # 按重要性排序,取前3个特征 sorted_importance = sorted(importance_dict.items(), key=lambda x: x[1], reverse=True)[:3] feature_names = [item[0] for item in sorted_importance] importance_scores = [item[1] for item in sorted_importance] # 创建画布和轴对象 fig, ax = plt.subplots(figsize=(10, 3)) # 绘制横向条形图 bars = ax.barh(range(len(feature_names)), importance_scores) # 设置y轴的特征标签 ax.set_yticks(range(len(feature_names))) ax.set_yticklabels(feature_names) # 给每个条形添加格式化后的数值标签 for bar in bars: bar_width = bar.get_width() # 把数值放在条形右侧居中位置,格式化为整数 ax.text(bar_width, bar.get_y() + bar.get_height()/2, f'{int(bar_width)}', ha='left', va='center', fontsize=10) # 添加坐标轴标题和图标题 ax.set_xlabel('Gain') ax.set_title('Top 3 Feature Importance') plt.show()
这种方式自由度拉满,你可以根据需求随意调整任何细节。
为啥你之前调整X轴格式没生效?
你说尝试X轴格式化没效果,其实是因为show_values=True显示的数值不是X轴的刻度标签,而是独立的文本元素,所以调整X轴刻度格式根本影响不到这些数值——这就是问题的核心啦~
内容的提问来源于stack exchange,提问作者Giedrė Oksaitė
相关产品推荐
相关产品推荐

