随机森林回归模型多次运行精度不同是否正常?该选用哪个精度?
Random Forest回归模型多次运行精度不一致的问题分析
每次从头运行Random Forest(RF)回归模型时,得到的精度都不同。运行代码如下:
df17_tmp1 = df17_tmp.sample(frac=6, replace = True).reset_index(drop=True) x_3d = df17_tmp1[col_in_3d] # Features; y_3d = df17_tmp1['over/under_exc_vol(m3)'].values # Target # In[29]: x_train_3d, x_test_3d, y_train_3d, y_test_3d = train_test_split(x_3d, y_3d, test_size = 0.3, random_state = 42) # # train RF # In[30]: x_train_3d = x_train_3d.fillna(0).reset_index(drop = True) x_test_3d = x_test_3d.fillna(0).reset_index(drop = True) y_train_3d[np.isnan(y_train_3d)] = 0 y_test_3d[np.isnan(y_test_3d)] = 0 rf_3d = RandomForestRegressor(n_estimators = 70, random_state = 42) rf_3d.fit(x_train_3d, y_train_3d) # # Predict with RF and evaluate # In[31]: prediction_3d = rf_3d.predict(x_test_3d) mse_3d = mean_squared_error(y_test_3d, prediction_3d) rmse_3d = mse_3d**.5 abs_diff_3d = np.array(np.abs((y_test_3d - prediction_3d)/y_test_3d)) abs_diff_3d = abs_diff_3d[~np.isinf(abs_diff_3d)] mape_3d = np.nanmean(abs_diff_3d)*100 accuracy_3d = 100 - mape_3d
多次运行得到的精度结果为:85.94、85.71、85.83、82.64、86.56、85.24、83.40、82.39、84.98、83.81。
问题
请问这种情况是否正常?应该选用哪个精度?
一、这种情况是否正常?
是正常的。核心原因是代码开头的重采样步骤未固定随机种子:
df17_tmp.sample(frac=6, replace=True)方法在未指定random_state时,每次运行会生成不同的采样数据集;- 尽管后续的
train_test_split和RandomForestRegressor都设置了random_state=42,但第一步采样的随机性会导致每次的训练/测试数据分布存在差异,最终影响模型的精度输出。
二、应该选用哪个精度?
- 优先采用统计值反映模型真实性能:
- 计算多次运行结果的平均值:你给出的10次结果平均值约为84.65,这个值能更客观体现模型的整体表现;
- 同时可以计算标准差(约1.5),它能反映模型性能的波动程度,帮助判断模型稳定性;
- 若需固定结果,消除随机性根源:
- 在
sample方法中添加random_state参数,比如:df17_tmp1 = df17_tmp.sample(frac=6, replace = True, random_state=42).reset_index(drop=True) - 这样每次运行的采样数据完全一致,后续模型的精度结果也会固定。
- 在
内容的提问来源于stack exchange,提问作者Abboud
相关产品推荐
相关产品推荐

