使用GridSearchCV调参时RandomForestRegressor预测偏差的优化参数问询
使用GridSearchCV时RandomForestRegressor的预测偏差
问题
针对GridSearchCV与RandomForestRegressor,移除或调整哪个参数可优化预测结果、提升模型表现?
代码实现
start = time.time() param_dists = {'max_features': ["sqrt", "log2", None], #默认值None 'max_depth': [None, range(1, 50)], #默认值None 'criterion': ['mae', 'squared_error'] #默认值squared_error } model_RandomForestRegressor_TuneSearchCV = RandomForestRegressor(random_state=270223) grid_RandomForestRegressor = GridSearchCV(model_RandomForestRegressor_TuneSearchCV, param_dists, cv=3, n_jobs=-1, scoring='neg_mean_absolute_error') grid_RandomForestRegressor.fit(X_train, y_train) best_score = -1 * grid_RandomForestRegressor.best_score_ predict_dt = grid_RandomForestRegressor.predict(X_test) time_hyperopt_RandomForestRegressor = round((time.time() - start), 2) print("调参后随机森林模型的MAE:", best_score) print('最佳参数', grid_RandomForestRegressor.best_params_) results_model = {'模型': '带调参的RandomForestRegressor', 'MAE': best_score, '总耗时':time_hyperopt_RandomForestRegressor } results = results.append(results_model, ignore_index=True)
当前模型情况
- 评估指标:MAE(平均绝对误差)
- 调参后MAE:5.564043154396905
- 最佳参数:
{'criterion': 'mae', 'max_depth': None, 'max_features': None}
参数优化建议
1. 聚焦调整max_depth参数
当前最佳参数中max_depth设为None,意味着决策树会完全生长至叶子节点纯净,极易引发过拟合。建议缩小深度搜索范围(比如5-20,而非1-50的宽泛区间),通过GridSearchCV测试不同深度对MAE的影响,找到拟合与泛化的平衡点。
2. 优化max_features的搜索维度
现有候选值仅包含sqrt、log2和None,可新增具体比例值(如特征总数的1/3、2/3)或小数比例(0.2、0.5、0.8),更精准地找到适配当前数据集的特征采样策略。
3. 新增关键参数至调参网格
现有参数覆盖范围有限,建议加入以下核心参数:
n_estimators:决策树数量,默认100,可测试50、200、500等值,更多树能提升模型稳定性,但需权衡计算成本。min_samples_split:拆分内部节点的最小样本数,默认2,尝试5、10、20等值,降低过拟合风险。min_samples_leaf:叶子节点的最小样本数,默认1,尝试2、5、10等值,限制树的生长复杂度。
4. 调整交叉验证策略
当前用cv=3,可增大折数至5或10,让模型评估更稳定,避免因数据划分偏差导致的参数选择误差。
5. 移除冗余参数测试
当前最佳参数多为默认值,说明现有网格未覆盖到更优取值。可先移除criterion参数(已选与评估指标一致的mae),专注优化max_depth和max_features的取值范围,或新增上述参数拓展搜索空间。
内容的提问来源于stack exchange,提问作者Kirill
相关产品推荐
相关产品推荐

