You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R中cv.glmnet针对二分类数据返回的MSE是否为实际值两倍?

二分类数据下cv.glmnet返回的MSE/MAE与实际计算值的两倍差异问题

当处理二分类数据时,R中cv.glmnet通过$cvm或plot()返回的MSE(均方误差)和MAE(平均绝对误差)最小值是实际计算值的两倍,这一差异造成了极大困扰。曾基于类似下方代码生成的图表投稿,图中y轴MAE值大于0.5(二分类场景下随机猜测的MAE约为0.5),被审稿人质疑模型表现差于随机猜测,但无法合理解释该差异。实际用预测值计算的MAE恰好是cv.glmnet返回值的一半。

以下是验证该差异的R代码:

# sample R code to illustrate huge discrepancy between MSE and MAE
# values from cv.glmnet when compared to prediction errors
# calculated applying a test data set 

# generate binomial data
n <- 10000
set.seed(1234)
x1 <- runif(n, -2, 2); x2 <- runif(n, -2, 2)
p <- exp(x1 + x2)/(1 + exp(x1 + x2))
y <- rbinom(n, 1, p)
  # Note: if the line above is replaced by  y <- 5*p + rnorm(n)
  #  and later family="gaussian" in cv.glmnet() 
  #  then there are no similar discrepancies as with binary data

# predictors
x <- matrix(0, n, 10)
for (k in 1:ncol(x)) {a <- runif(1) 
 x[, k] <- 0.5*(a*x1 + (1-a)*x2 + runif(n, -2, 2))}

# training data
xa <- x[seq(n/2), ]; ya <- y[seq(n/2)]
# test data
xb <- x[-1*seq(n/2), ]; yb <- y[-1*seq(n/2)]

install.packages("glmnet")
library(glmnet)
# set alpha=1 i.e. apply Lasso regression with optimal regularization
# parameter lambda chosen by MSE or MAE criterion
cvfit_MSE <- cv.glmnet(xa, ya, family="binomial", alpha=1, 
                       type.measure="mse", nfolds=10)
cvfit_MAE <- cv.glmnet(xa, ya, family="binomial", alpha=1, 
                       type.measure="mae", nfolds=10)

min(cvfit_MSE$cvm) # MSE=0.3734 (mean of squared errors)
min(cvfit_MAE$cvm) # MAE=0.7477 (mean of absolute values of the errors)
plot(cvfit_MSE) 
           # also this has min 0.3734 on y-axis (labelled as "MSE")
plot(cvfit_MAE) # here y-axis (labelled as MAE) starts at 0.75 
                # but even random guessing would produce MAE about 0.5, 
                # extremely confusing... should y-axis label be 2*MAE??

# obtain predictions for the test data:
pp <- predict(cvfit_MSE, newx=xb, s="lambda.min", type="response")

# compare these to the actual observed values (variable "yb") in test data.
# surprisingly, MSE and MAE from these predictions are half of those from cv.glmnet:
mean((pp-yb)^2) # MSE=0.17650 (mean of squared errors)
mean(abs(pp-yb)) # MAE=0.36380 (mean of absolute values of the errors)

运行代码后可观察到明显差异:

  • cv.glmnet返回的最小MSE为0.3734,最小MAE为0.7477
  • 用测试集预测值手动计算的实际MSE为0.17650,实际MAE为0.36380,恰好是前者的一半

内容的提问来源于stack exchange,提问作者Mark Nh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 15:57:47