Sklearn随机森林:获取网格搜索选定的特征名称
随机森林max_features相关问题解答
你的代码示例:
X_train, X_test, y_train, y_test = train_test_split(features, target, test_size=0.15, random_state=42) model = RandomForestRegressor() param_grid = { 'n_estimators': [100, 200, 500],#, 300], 'max_depth': [3, 5, 7], 'max_features': [3, 5, 7], 'random_state': [42] }
网格搜索代码:
grid_search = GridSearchCV(model, param_grid, cv=5) grid_search.fit(X_train, y_train) print(grid_search.best_params_)
最优参数输出:
{'max_depth': 7, 'max_features': 3, 'n_estimators': 500, 'random_state': 42}
预测及R2计算代码:
y_train_pred = best_model.predict(X_train) y_test_pred = best_model.predict(X_test) train_r2 = r2_score(y_train, y_train_pred) test_r2 = r2_score(y_test, y_test_pred)
问题解答
不是仅用3个特征计算R2
max_features=3的含义是:随机森林里每棵决策树在分裂节点时,会从所有特征中随机挑选3个作为候选,用来寻找最优分裂点。整个模型训练时依然用到了全部输入特征,不同树挑选的3个特征可能不同,最终预测结果是所有树的预测平均值。你得到的R2是基于全特征训练的随机森林模型计算的,并非只依赖3个特征。获取对模型贡献最大的前3个特征名称
如果要找出对模型预测最关键的3个特征,可以借助随机森林的feature_importances_属性,它会输出每个特征的重要性得分。代码示例如下:
# 获取网格搜索得到的最优模型 best_model = grid_search.best_estimator_ # 获取特征名称(假设features是DataFrame,直接取列名;若为数组需提前保存特征名列表) feature_names = features.columns # 关联特征名称与对应重要性得分 feature_importance = list(zip(feature_names, best_model.feature_importances_)) # 按重要性降序排序 feature_importance_sorted = sorted(feature_importance, key=lambda x: x[1], reverse=True) # 提取并输出前3个最重要的特征 top_3_features = [item[0] for item in feature_importance_sorted[:3]] print("对模型贡献最大的前3个特征:", top_3_features)
内容的提问来源于stack exchange,提问作者Sinha
相关产品推荐
相关产品推荐

