You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于4折交叉验证计算随机森林回归模型的平均R2 Score、RMSE与MAPE

没问题,我来帮你调整代码实现4折交叉验证,并计算这三个评估指标的平均值。下面是修改后的完整代码,以及关键部分的说明:

修改后的完整代码

# STEP 1: 导入所需库与模块
import pandas as pd
import numpy as np
from sklearn.tree import DecisionTreeRegressor
from sklearn.metrics import r2_score
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_squared_error
from sklearn import preprocessing
from sklearn.metrics import mean_absolute_percentage_error
from sklearn.ensemble import RandomForestRegressor
from sklearn.model_selection import KFold
import matplotlib.pyplot as plt
import seaborn as sns
from sklearn.svm import SVR
from sklearn.feature_selection import RFE, SelectFromModel
from sklearn.model_selection import cross_val_score
from sklearn import model_selection
import math

# STEP 2: 读取数据并执行基础数据检查
path = "C:/AKHIL/OTHER/MISSION_INTERNSHIP/Space4Good/DATA_FOR_MODELING_2.xlsx"
sheet_1 = pd.read_excel(path, sheet_name='Model Development')
sheet_2 = pd.read_excel(path, sheet_name='Validation Data')
x_validation = sheet_2.drop(['ID'], axis=1).values
print("STATISTICAL DESCRIPTION:")
print(sheet_1.describe(), "\n")

# STEP 3: 创建特征与响应变量数组
target_column = ['y', 'ID']
predictors = list(set(list(sheet_1.columns)) - set(target_column))

# STEP 4: 通过缩放将预测变量归一化至0-1区间
scaler = preprocessing.MinMaxScaler(feature_range=(0, 1))
names = sheet_1[predictors].columns
d = scaler.fit_transform(sheet_1[predictors])
sheet_1[predictors] = pd.DataFrame(d, columns=names)
print("STATISTICAL DESCRIPTION AFTER NORMALIZATION:")
print(sheet_1.describe())

# STEP 5: 准备完整的特征和目标变量(不再用train_test_split,改用K-Fold)
X = sheet_1[predictors] # 所有值已归一化
y = sheet_1['y']

# STEP 6 & 7: 4折交叉验证 + 模型评估
# 初始化列表存储每折的评估结果
r2_scores = []
rmse_scores = []
mape_scores = []

# 定义4折交叉验证,加上shuffle确保数据分布均匀
kf = KFold(n_splits=4, shuffle=True, random_state=0)

for fold, (train_index, test_index) in enumerate(kf.split(X), 1):
    print(f"=== Fold {fold} ===")
    # 划分当前折的训练集和测试集
    X_train, X_test = X.iloc[train_index], X.iloc[test_index]
    y_train, y_test = y.iloc[train_index], y.iloc[test_index]
    
    # 初始化并训练随机森林模型(每次折都重新初始化,避免模型状态残留)
    dtree = RandomForestRegressor(n_estimators=500, oob_score=True, random_state=0)
    dtree.fit(X_train, y_train)
    
    # 预测
    pred_test_tree = dtree.predict(X_test)
    
    # 计算当前折的评估指标
    r2 = r2_score(y_test, pred_test_tree)
    rmse = np.sqrt(mean_squared_error(y_test, pred_test_tree))
    mape = mean_absolute_percentage_error(y_test, pred_test_tree)
    
    # 保存当前折的结果
    r2_scores.append(r2)
    rmse_scores.append(rmse)
    mape_scores.append(mape)
    
    # 打印当前折的结果
    print(f"RMSE: {rmse:.4f}")
    print(f"R2 SCORE: {r2:.4f}")
    print(f"MAPE: {mape:.4f}\n")

# 计算并打印所有折的平均指标
print("=== 4折交叉验证平均结果 ===")
print(f"平均R2 SCORE: {np.mean(r2_scores):.4f} (标准差: {np.std(r2_scores):.4f})")
print(f"平均RMSE: {np.mean(rmse_scores):.4f} (标准差: {np.std(rmse_scores):.4f})")
print(f"平均MAPE: {np.mean(mape_scores):.4f} (标准差: {np.std(mape_scores):.4f})")

关键修改说明

  1. 移除固定train-test split:K-Fold会遍历整个数据集做验证,不需要单独划分固定的训练/测试集
  2. 添加shuffle=True:默认KFold不打乱数据,加上这个参数能让每折的数据分布更均匀,避免因数据排序导致的偏差
  3. 指标存储列表:创建三个列表保存每折的R2、RMSE、MAPE结果,方便后续计算平均值
  4. 每折重新初始化模型:每次循环都新建RandomForestRegressor实例,确保模型不受上一折训练的状态影响
  5. 输出平均值+标准差:除了平均指标,标准差能帮你了解模型在不同数据子集上的性能稳定性

这样你就能得到更可靠的模型评估结果,比单一的train-test split更能反映模型的泛化能力!

内容的提问来源于stack exchange,提问作者Akhil Chibber

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 13:47:47