You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何同一MLP模型采用不同损失计算方式得到的误差值存在差异?

多层感知器不同误差计算方式结果差异的原因及选择建议

问题背景

用户使用MLPClassifier实现分类模型,通过GridSearchCV调参后,用多种方式计算训练/测试误差,得到了差异明显的结果:

调参与训练代码

hl_parameters = {'hidden_layer_sizes': [(10,), (50,), (10,10,), (50,50,)]}

mlp_cv = MLPClassifier(max_iter=300, alpha=1e-4, solver='sgd', tol=1e-4, learning_rate_init=.1, verbose=True, random_state=ID)
mlp_cv.fit(X_train, y_train)

clf = GridSearchCV(mlp_cv, hl_parameters, cv=5)
clf.fit(X_train, y_train)

print(clf.best_params_)
print(clf.best_score_)

调参输出:

Best parameters set found: {'hidden_layer_sizes': (50,)} 
Score with best parameters:
0.7779999999999999

误差计算代码

mlp = clf.best_estimator_

y_trainpred= clf.predict(X_train)
training_error = np.mean(np.square(np.array(y_trainpred)-np.array(y_train)))
y_testpred= clf.predict(X_test)
test_error = np.mean(np.square(np.array(y_testpred)-np.array(y_test)))

print ("Best NN training error: %f" % training_error)
print('Best NN training error2: ', 1. - clf.best_score_)
print('Best NN training error3: ', 1. - mlp.score(X_train,y_train))
print('Best NN test error2: ', 1. - mlp.score(X_test,y_test))
print ("Best NN test error: %f" % test_error)

误差输出:

Best NN training error: 0.000000
Best NN training error2:  0.2220000000000001
Best NN training error3:  0.0
Best NN test error2:  0.21289075630252097
Best NN test error: 2.588303

误差差异的核心原因

1. 指标类型不匹配

你用np.mean(np.square(...))计算的是均方误差(MSE),这是回归任务的标准指标,但你的模型是MLPClassifier(分类模型)。分类任务中,标签通常是离散类别(比如0/1/2...),用平方差衡量完全没有意义:

  • 训练集MSE为0,只是因为模型在训练集上预测完全正确,差值平方后平均为0;
  • 测试集MSE为2.58,是因为部分样本预测错误,标签差值的平方被放大,无法反映分类的真实错误程度。

2. 计算的数据集不同

  • clf.best_score_是GridSearchCV在5折交叉验证的验证集上的平均准确率,所以1 - clf.best_score_是交叉验证验证集的平均错误率,不是训练集的错误率;
  • mlp.score(X_train,y_train)是最优模型在完整训练集上的准确率,1 - 该值是训练集的分类错误率,和交叉验证的验证集错误率自然存在差异。

3. 模型训练目标混淆

MLPClassifier的训练目标是最小化交叉熵损失(分类任务的标准损失),而非均方误差。用MSE评估分类模型,本质是用回归指标衡量分类任务,逻辑上不成立,结果没有参考价值。

应该选用哪种结果?

对于分类任务,只需要关注基于准确率的错误率:

  • 训练集错误率:选1 - mlp.score(X_train,y_train),它反映模型在训练数据上的拟合程度,你的结果为0说明模型完全拟合训练集;
  • 测试集错误率:选1 - mlp.score(X_test,y_test),这是模型在 unseen 数据上的泛化能力指标,你的结果约0.21,说明测试集上约21%的样本预测错误,是你需要重点关注的核心指标;
  • clf.best_score_对应的错误率(0.22)是交叉验证的平均验证错误率,可用来辅助判断模型泛化稳定性,和测试集错误率接近,说明模型表现稳定。

完全放弃用均方误差评估分类模型,这是回归任务的专属指标,不适合分类场景。

内容的提问来源于stack exchange,提问作者cipos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 16:10:31