You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中随机森林分类器模型性能评估代码求助

帮你修正随机森林分类器的性能评估代码

嘿,我来帮你搞定这个随机森林调参和性能评估的问题!你的代码踩了几个关键的坑,导致没法有效评估模型的真实泛化能力,我来给你拆解并修正:

你的代码存在的核心问题

  • 只在训练集上计算得分:这样得到的是模型在“见过的数据”上的准确率,会严重高估模型性能,完全没法反映它在新数据上的表现。
  • 分数存储逻辑错误:你把所有参数组合的得分都塞进一个一维列表,最后用np.argmax(score)找最大值时,根本没法对应回正确的max_depth和n_estimators组合——因为列表的索引是按循环顺序依次累加的,和两个参数的二维组合不匹配。

推荐的修正方案:用GridSearchCV自动调参(简洁高效)

sklearn自带的GridSearchCV可以帮你自动遍历所有参数组合,用交叉验证评估每个组合的泛化能力,还能直接给出最优参数,省去手动处理索引的麻烦:

from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import GridSearchCV
import numpy as np

# 定义要搜索的参数网格
param_grid = {
    'max_depth': np.arange(1, 30),
    'n_estimators': np.arange(1, 100)
}

# 初始化随机森林分类器,加random_state保证实验可复现
rf_model = RandomForestClassifier(random_state=42)

# 初始化网格搜索:用5折交叉验证,评估指标选准确率
grid_search = GridSearchCV(
    estimator=rf_model,
    param_grid=param_grid,
    cv=5,  # 5折交叉验证
    scoring='accuracy',
    n_jobs=-1  # 用所有CPU核心加速,可选
)

# 在训练集上拟合搜索最优参数
grid_search.fit(X_train, y_train)

# 输出结果
print(f"最优参数组合: {grid_search.best_params_}")
print(f"最优交叉验证准确率: {grid_search.best_score_:.4f}")

# 用最优参数在测试集上评估最终性能
best_rf = grid_search.best_estimator_
test_accuracy = best_rf.score(X_test, y_test)
print(f"测试集最终准确率: {test_accuracy:.4f}")

如果你想手动实现循环(更灵活)

要是你偏好手动控制循环逻辑,可以这样改进,核心是用验证集评估+记录参数与得分的对应关系:

from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split
import numpy as np

# 先把原始训练集拆分为训练集+验证集(用来评估参数)
X_train_split, X_val, y_train_split, y_val = train_test_split(
    X_train, y_train, test_size=0.2, random_state=42
)

depth_range = np.arange(1, 30)
estimators_range = np.arange(1, 100)

# 存储每个参数组合的得分,以及跟踪最优结果
best_score = 0
best_params = {}
param_score_map = {}

for depth in depth_range:
    for n_est in estimators_range:
        # 初始化模型,固定random_state保证可复现
        rf = RandomForestClassifier(n_estimators=n_est, max_depth=depth, random_state=42)
        rf.fit(X_train_split, y_train_split)
        
        # 用验证集计算得分,而不是训练集!
        val_score = rf.score(X_val, y_val)
        param_score_map[(depth, n_est)] = val_score
        
        # 更新最优参数
        if val_score > best_score:
            best_score = val_score
            best_params = {'max_depth': depth, 'n_estimators': n_est}

# 输出最优参数和验证集得分
print(f"最优参数组合: {best_params}")
print(f"最优验证集准确率: {best_score:.4f}")

# 最后用最优参数在完整训练集上重新训练,再评估测试集
final_rf = RandomForestClassifier(**best_params, random_state=42)
final_rf.fit(X_train, y_train)
test_score = final_rf.score(X_test, y_test)
print(f"测试集最终准确率: {test_score:.4f}")

额外提醒

  • 一定要用验证集或交叉验证评估模型,绝对不能只看训练集得分,否则会选到过拟合严重的参数。
  • 加上random_state参数,保证每次运行的结果一致,方便调试。
  • 如果你的参数范围太大(比如n_estimators到100),可以先做“粗调”(比如步长设为5,np.arange(10,100,5)),找到大致最优范围后再缩小范围做“细调”,能大幅节省时间。

内容的提问来源于stack exchange,提问作者Lulu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:20:21