You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Sklearn随机森林:获取网格搜索选定的特征名称

随机森林max_features相关问题解答

你的代码示例:

X_train, X_test, y_train, y_test = train_test_split(features, target, test_size=0.15, random_state=42)

model = RandomForestRegressor()
param_grid = {
    'n_estimators': [100, 200, 500],#, 300],
    'max_depth': [3, 5, 7],
    'max_features': [3, 5, 7],
    'random_state': [42]
}

网格搜索代码:

grid_search = GridSearchCV(model, param_grid, cv=5)
grid_search.fit(X_train, y_train)
print(grid_search.best_params_)

最优参数输出:

{'max_depth': 7, 'max_features': 3, 'n_estimators': 500, 'random_state': 42}

预测及R2计算代码:

y_train_pred = best_model.predict(X_train)
y_test_pred = best_model.predict(X_test)
train_r2 = r2_score(y_train, y_train_pred)
test_r2 = r2_score(y_test, y_test_pred)

问题解答

  1. 不是仅用3个特征计算R2
    max_features=3的含义是:随机森林里每棵决策树在分裂节点时,会从所有特征中随机挑选3个作为候选,用来寻找最优分裂点。整个模型训练时依然用到了全部输入特征,不同树挑选的3个特征可能不同,最终预测结果是所有树的预测平均值。你得到的R2是基于全特征训练的随机森林模型计算的,并非只依赖3个特征。

  2. 获取对模型贡献最大的前3个特征名称
    如果要找出对模型预测最关键的3个特征,可以借助随机森林的feature_importances_属性,它会输出每个特征的重要性得分。代码示例如下:

# 获取网格搜索得到的最优模型
best_model = grid_search.best_estimator_

# 获取特征名称(假设features是DataFrame,直接取列名;若为数组需提前保存特征名列表)
feature_names = features.columns

# 关联特征名称与对应重要性得分
feature_importance = list(zip(feature_names, best_model.feature_importances_))

# 按重要性降序排序
feature_importance_sorted = sorted(feature_importance, key=lambda x: x[1], reverse=True)

# 提取并输出前3个最重要的特征
top_3_features = [item[0] for item in feature_importance_sorted[:3]]
print("对模型贡献最大的前3个特征:", top_3_features)

内容的提问来源于stack exchange,提问作者Sinha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 08:43:15