GAM预测器结构:矩阵与长数据框格式拟合结果差异问询
带滞后结构的GAM拟合:宽格式与长格式数据的结果差异问题
模型与数据背景
我正在拟合一个典型的带滞后结构的GAM,模型基本形式为:
g1 <- gam(y~te(e, l), data=df)
数据生成代码
# 暴露变量 e <- matrix(rnorm(2500, mean = 100, sd = 25), ncol=50) # 滞后变量 l <- matrix(rep(seq(49,0),50), ncol = 50, byrow=T) # 响应变量y:暴露e按滞后权重求和(GAM需要拟合的目标) we <- matrix(sample(seq(0.01, 0.5, 0.02),2500, replace = T), ncol = 50, byrow=T) # 随机权重 ex <- e*we # 按滞后加权后的暴露 df <- data.frame(y=apply(ex, 1, FUN = sum)) # 响应变量
上述数据组织方式遵循S.Wood所著《Generalized Additive Models: an introduction in R》中的建议。
宽格式数据的GAM拟合结果
Family: gaussian Link function: identity Formula: y ~ te(e, l) Parametric coefficients: Estimate Std. Error t value Pr(>|t|) (Intercept) 1267.97 12.63 100.4 <2e-16 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 Approximate significance of smooth terms: edf Ref.df F p-value te(e,l) 2.605 3.054 4.808 0.00512 ** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 Rank: 24/25 R-sq.(adj) = 0.208 Deviance explained = 25% GCV = 8595 Scale est. = 7975.3 n = 50
长格式数据的转换与拟合结果
将数据转换为长格式:
df <- data.frame(l=rep(seq(0,49),each=50), e=c(e[,seq(50,1)]), y=rep(y$y, 50))
对应的GAM输出:
Formula: y ~ te(e, l) Parametric coefficients: Estimate Std. Error t value Pr(>|t|) (Intercept) 1267.966 1.984 639.2 <2e-16 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 Approximate significance of smooth terms: edf Ref.df F p-value te(e,l) 4.11 4.848 2.215 0.0448 * --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 R-sq.(adj) = 0.00337 Deviance explained = 0.501% GCV = 9858.2 Scale est. = 9838.1 n = 2500
核心疑问
无论当前拟合结果好坏,我想知道为什么将数据组织为长数据框格式后,GAM的拟合结果会出现如此明显的差异?
若问题较为基础,深表歉意,但我非常希望得到相关反馈及参考资料以更好地理解这一点。谢谢。
内容的提问来源于stack exchange,提问作者monviso
相关产品推荐
相关产品推荐

