为何GLM模型预测值未呈现曲线,而是三段直线?
问题描述
我编写了一段脚本,用于构建GLM模型并将数据集归一化至0-1区间,随后绘图展示变量间关系。处理多个数据集时,绘图均为曲线(如下左图),但针对某一特定数据集,绘图却呈现三段直线(如下右图)。我猜测这与predict函数中的newdata参数有关,但无法确定。

原代码
# turn off scientific notation options(scipen = 999) # recreating the data IV_BP <- structure(list(Breakpoints = c("Min", "BP1", "BP2", "BP3", "BP4", "Max"), SES = c(-1.8, -0.3, -0.1, 0.1, 0.3, 0.8), Normalised_value = c(0,0.2, 0.4, 0.6, 0.8, 1)), class = "data.frame", row.names = c(NA, -6L)) IV_df <- structure(list(SES = c(-0.006, 0.078, 0.028, -0.066, 0.041, -0.025, 0.006, -0.021, -0.013, -0.145, -0.065, 0.026, 0.068, -0.22, 0.138, 0.019, 0.174, 0.107, 0.339, 0.219, 0.093, -0.057, -0.19, 0.01, 0.085, -0.011, -0.075, -0.113, -0.019, 0.141, -0.045, -0.258, -0.02, -0.178, -0.142, -0.067, 0.1, -0.155, 0.007, -0.18, -0.258, -0.497)), class = "data.frame", row.names = c(NA, -42L)) # make glm glmfit <- glm(Normalised_value~SES,data=IV_BP,family = quasibinomial()) # use glm to transform values IV_df$CC_Transformed <- predict(glmfit,newdata=IV_df,type="response") # make a graph plot(IV_BP$SES, IV_BP$Normalised_value, xlab = "Socioeconomic Status Index Score", ylab = "Normalised Values", xlim = c(-2, 2), pch = 19, col = "blue", panel.first = c(abline(h = 0, col = "lightgrey"), abline(h = 0.2, col = "lightgrey"), abline(h = 0.4, col = "lightgrey"), abline(h = 0.6, col = "lightgrey"), abline(h = 0.8, col = "lightgrey"), abline(h = 1, col = "lightgrey"), lines(-2:2,predict(glmfit,newdata=data.frame(SES=-2:2),type="response"), col = "lightblue", lwd = 5)))
问题原因与解决办法
问题根本不是newdata参数的问题,而是绘图时生成预测点用的-2:2是整数序列,仅包含-2, -1, 0, 1, 2这5个离散点。用这么少的点连线,必然会呈现分段直线,而非平滑曲线。
要得到平滑的曲线,需要生成足够密集的连续x值。可以用seq(-2, 2, by = 0.01)生成从-2到2、步长0.01的密集序列,这样预测出的点数量足够多,连线后就是自然的平滑曲线。
修改后的代码
# 关闭科学计数法 options(scipen = 999) # 重建数据 IV_BP <- structure(list(Breakpoints = c("Min", "BP1", "BP2", "BP3", "BP4", "Max"), SES = c(-1.8, -0.3, -0.1, 0.1, 0.3, 0.8), Normalised_value = c(0,0.2, 0.4, 0.6, 0.8, 1)), class = "data.frame", row.names = c(NA, -6L)) IV_df <- structure(list(SES = c(-0.006, 0.078, 0.028, -0.066, 0.041, -0.025, 0.006, -0.021, -0.013, -0.145, -0.065, 0.026, 0.068, -0.22, 0.138, 0.019, 0.174, 0.107, 0.339, 0.219, 0.093, -0.057, -0.19, 0.01, 0.085, -0.011, -0.075, -0.113, -0.019, 0.141, -0.045, -0.258, -0.02, -0.178, -0.142, -0.067, 0.1, -0.155, 0.007, -0.18, -0.258, -0.497)), class = "data.frame", row.names = c(NA, -42L)) # 构建GLM模型 glmfit <- glm(Normalised_value~SES,data=IV_BP,family = quasibinomial()) # 用模型转换数值 IV_df$CC_Transformed <- predict(glmfit,newdata=IV_df,type="response") # 绘图:生成密集x序列保证曲线平滑 x_seq <- seq(-2, 2, by = 0.01) plot(IV_BP$SES, IV_BP$Normalised_value, xlab = "社会经济地位指数得分", ylab = "归一化值", xlim = c(-2, 2), pch = 19, col = "blue", panel.first = c(abline(h = 0, col = "lightgrey"), abline(h = 0.2, col = "lightgrey"), abline(h = 0.4, col = "lightgrey"), abline(h = 0.6, col = "lightgrey"), abline(h = 0.8, col = "lightgrey"), abline(h = 1, col = "lightgrey"), lines(x_seq,predict(glmfit,newdata=data.frame(SES=x_seq),type="response"), col = "lightblue", lwd = 5)))
内容的提问来源于stack exchange,提问作者Xandian97
相关产品推荐
相关产品推荐

