Python中随机森林分类器模型性能评估代码求助
帮你修正随机森林分类器的性能评估代码
嘿,我来帮你搞定这个随机森林调参和性能评估的问题!你的代码踩了几个关键的坑,导致没法有效评估模型的真实泛化能力,我来给你拆解并修正:
你的代码存在的核心问题
- 只在训练集上计算得分:这样得到的是模型在“见过的数据”上的准确率,会严重高估模型性能,完全没法反映它在新数据上的表现。
- 分数存储逻辑错误:你把所有参数组合的得分都塞进一个一维列表,最后用
np.argmax(score)找最大值时,根本没法对应回正确的max_depth和n_estimators组合——因为列表的索引是按循环顺序依次累加的,和两个参数的二维组合不匹配。
推荐的修正方案:用GridSearchCV自动调参(简洁高效)
sklearn自带的GridSearchCV可以帮你自动遍历所有参数组合,用交叉验证评估每个组合的泛化能力,还能直接给出最优参数,省去手动处理索引的麻烦:
from sklearn.ensemble import RandomForestClassifier from sklearn.model_selection import GridSearchCV import numpy as np # 定义要搜索的参数网格 param_grid = { 'max_depth': np.arange(1, 30), 'n_estimators': np.arange(1, 100) } # 初始化随机森林分类器,加random_state保证实验可复现 rf_model = RandomForestClassifier(random_state=42) # 初始化网格搜索:用5折交叉验证,评估指标选准确率 grid_search = GridSearchCV( estimator=rf_model, param_grid=param_grid, cv=5, # 5折交叉验证 scoring='accuracy', n_jobs=-1 # 用所有CPU核心加速,可选 ) # 在训练集上拟合搜索最优参数 grid_search.fit(X_train, y_train) # 输出结果 print(f"最优参数组合: {grid_search.best_params_}") print(f"最优交叉验证准确率: {grid_search.best_score_:.4f}") # 用最优参数在测试集上评估最终性能 best_rf = grid_search.best_estimator_ test_accuracy = best_rf.score(X_test, y_test) print(f"测试集最终准确率: {test_accuracy:.4f}")
如果你想手动实现循环(更灵活)
要是你偏好手动控制循环逻辑,可以这样改进,核心是用验证集评估+记录参数与得分的对应关系:
from sklearn.ensemble import RandomForestClassifier from sklearn.model_selection import train_test_split import numpy as np # 先把原始训练集拆分为训练集+验证集(用来评估参数) X_train_split, X_val, y_train_split, y_val = train_test_split( X_train, y_train, test_size=0.2, random_state=42 ) depth_range = np.arange(1, 30) estimators_range = np.arange(1, 100) # 存储每个参数组合的得分,以及跟踪最优结果 best_score = 0 best_params = {} param_score_map = {} for depth in depth_range: for n_est in estimators_range: # 初始化模型,固定random_state保证可复现 rf = RandomForestClassifier(n_estimators=n_est, max_depth=depth, random_state=42) rf.fit(X_train_split, y_train_split) # 用验证集计算得分,而不是训练集! val_score = rf.score(X_val, y_val) param_score_map[(depth, n_est)] = val_score # 更新最优参数 if val_score > best_score: best_score = val_score best_params = {'max_depth': depth, 'n_estimators': n_est} # 输出最优参数和验证集得分 print(f"最优参数组合: {best_params}") print(f"最优验证集准确率: {best_score:.4f}") # 最后用最优参数在完整训练集上重新训练,再评估测试集 final_rf = RandomForestClassifier(**best_params, random_state=42) final_rf.fit(X_train, y_train) test_score = final_rf.score(X_test, y_test) print(f"测试集最终准确率: {test_score:.4f}")
额外提醒
- 一定要用验证集或交叉验证评估模型,绝对不能只看训练集得分,否则会选到过拟合严重的参数。
- 加上
random_state参数,保证每次运行的结果一致,方便调试。 - 如果你的参数范围太大(比如n_estimators到100),可以先做“粗调”(比如步长设为5,
np.arange(10,100,5)),找到大致最优范围后再缩小范围做“细调”,能大幅节省时间。
内容的提问来源于stack exchange,提问作者Lulu
相关产品推荐
相关产品推荐

