运行DecisionTreeRegressor后如何选择特征并提取Top20-30高重要性特征
特征筛选实现方案
你可以直接在原有代码基础上追加以下逻辑实现需求,两种常用筛选方式如下:
方案1:固定取前N个重要特征(可自定义20-30区间取值)
import numpy as np # 将特征索引和重要性绑定后按得分降序排序 feature_importance = sorted(enumerate(importance), key=lambda x: x[1], reverse=True) # 自定义要取的特征数量,比如取前25个,可在20-30之间自行调整 top_k = 25 # 提取前N个特征的索引 top_k_idx = [item[0] for item in feature_importance[:top_k]] # 打印验证筛选结果 print(f"前{top_k}个重要特征索引及得分:") for idx in top_k_idx: print(f"Feature: {idx}, Score: {importance[idx]:.5f}") # 赋值筛选后的特征到新变量(数组格式数据集写法) X_train_selected = X_train_total[:, top_k_idx] # 如果你的数据集是Pandas的DataFrame格式,用下面的写法: # X_train_selected = X_train_total.iloc[:, top_k_idx]
方案2:按重要性阈值筛选
你提供的运行输出中大量特征得分为0或接近0,也可以通过设置阈值过滤无效特征:
# 自定义阈值,比如保留得分大于0.001的所有特征 threshold = 0.001 selected_idx = [i for i, score in enumerate(importance) if score > threshold] print(f"阈值筛选后有效特征数量:{len(selected_idx)}") X_train_selected = X_train_total[:, selected_idx]
补充说明
如果你的原始特征有自定义列名(数据集为Pandas DataFrame格式),可以通过索引映射拿到对应列名,更便于后续特征分析:
feature_names = X_train_total.columns.tolist() selected_feature_names = [feature_names[i] for i in top_k_idx]
内容的提问来源于stack exchange,提问作者Dataleon
相关产品推荐
相关产品推荐

