如何过滤低重要性特征 减少随机森林特征重要性柱状图柱数
问题解答
过滤随机森林特征重要性柱状图中低重要性条目、解决图表杂乱的需求完全可以实现,以下提供两种可直接复用的实现方案,可根据你的数据集特征规模选择。
你当前生成的杂乱柱状图效果如下:
原有实现代码:
rf = RandomForestRegressor(n_estimators=100, max_depth=3) rf.fit(X_train, y_train) sorted_idx = rf.feature_importances_.argsort() plt.figure(figsize=(8, 30)) plt.barh(X_train.columns[sorted_idx], rf.feature_importances_[sorted_idx]) plt.xlabel("Random Forest Feature Importance")
方案1:按重要性数值阈值过滤
自定义重要性阈值(可根据特征重要性分布调整,常用值为0.005、0.01),仅保留重要性高于阈值的特征绘图:
import numpy as np rf = RandomForestRegressor(n_estimators=100, max_depth=3) rf.fit(X_train, y_train) # 可根据实际展示效果调整阈值大小 importance_threshold = 0.01 # 筛选达标的特征与对应重要性 mask = rf.feature_importances_ >= importance_threshold selected_features = X_train.columns[mask] selected_importances = rf.feature_importances_[mask] # 对筛选结果排序 sorted_idx = selected_importances.argsort() # 画布高度自适应筛选后的特征数量,避免留白过多 plt.figure(figsize=(8, len(selected_features)*0.4)) plt.barh(selected_features[sorted_idx], selected_importances[sorted_idx]) plt.xlabel("Random Forest Feature Importance") plt.tight_layout() # 自动调整布局避免标签截断
方案2:按Top N高重要性特征过滤(更推荐)
直接指定保留重要性排名前N位的特征,无需反复调试阈值,展示效果更可控:
import numpy as np rf = RandomForestRegressor(n_estimators=100, max_depth=3) rf.fit(X_train, y_train) # 自定义要展示的高重要性特征数量 top_n = 20 # 取重要性最高的前N个特征索引 sorted_idx = rf.feature_importances_.argsort()[-top_n:] plt.figure(figsize=(8, top_n*0.4)) plt.barh(X_train.columns[sorted_idx], rf.feature_importances_[sorted_idx]) plt.xlabel("Random Forest Feature Importance") plt.tight_layout()
补充提示:如果需要保留完整信息占比,可以额外增加一个「其他低重要性特征」柱形,将所有被过滤特征的重要性求和作为该柱形的数值,不会丢失整体重要性分布信息。
内容的提问来源于stack exchange,提问作者leskovecg
相关产品推荐
相关产品推荐

