mgcviz绘制拟合QGAM时散点图异常的原因及解决方法
问题:mgcviz绘制QGAM时散点与实际数据不符
问题背景与复现代码
首先,常规GAM的散点图与拟合结果可通过以下代码实现:
#### Libraries #### library(tidyverse) library(mgcViz) library(qgam) library(quantreg) #### Save Data as Tibble #### data("barro") tib <- as_tibble(MASS::mcycle) #### Inspect Scatterplot #### tib %>% ggplot(aes(x=times, y=accel))+ geom_point(alpha=.4)+ theme_classic()+ geom_smooth(method = "gam", formula = y ~ s(x))+ labs(x="Times", y="Acceleration", title = "Normal Fitted GAM")
在此基础上拟合多分位数QGAM的代码如下:
#### Set Quants and Fit Multiple Quantile GAM #### q <- c(.2,.5,.8) fit <- mqgam(accel ~ s(times), data = tib, q=q)
但使用mgcviz绘制分位数拟合图时,发现散点位置与原始数据不符:
#### Save QDO Objects #### q2 <- qdo(fit,.2) #### Grab Visual Data #### pg2 <- getViz(q2) #### Plot in MGCVIZ #### final.plot <- plot(pg2, select = 1) final.plot + l_fitLine(colour = "red") + l_rug(mapping = aes(x=x, y=y), alpha = 0.8) + l_ciLine(mul = 5, colour = "blue", linetype = 2) + l_points(shape = 19, size = 1, alpha = 0.1) + theme_classic()
上述代码生成的图中,部分y轴数据点位置与原散点图不一致。
问题原因
mgcviz在处理qgam对象时,plot()默认提取的是模型拟合的分位数响应值,而非原始数据集的y值。l_points()调用的是模型对象内部存储的拟合相关数据,并非原始观测点,因此出现位置偏差。
解决方法
方法1:手动添加原始散点(推荐)
直接在mgcviz的绘图对象上,使用ggplot2的geom_point()调用原始数据集,替换l_points():
final.plot <- plot(pg2, select = 1) final.plot + l_fitLine(colour = "red") + l_rug(mapping = aes(x=x, y=y), alpha = 0.8) + l_ciLine(mul = 5, colour = "blue", linetype = 2) + # 添加原始数据的散点 geom_point(data = tib, aes(x = times, y = accel), shape = 19, size = 1, alpha = 0.1) + theme_classic()
方法2:修改mgcviz对象的数据源(不推荐)
若坚持使用l_points(),可手动将原始数据注入到pg2对象的data字段中,但需修改对象内部结构,操作相对繁琐:
# 将原始数据加入pg2的data中,并重命名匹配绘图映射 pg2$data <- tib %>% rename(x = times, y = accel) # 重新绘图 final.plot <- plot(pg2, select = 1) final.plot + l_fitLine(colour = "red") + l_rug(mapping = aes(x=x, y=y), alpha = 0.8) + l_ciLine(mul = 5, colour = "blue", linetype = 2) + l_points(shape = 19, size = 1, alpha = 0.1) + theme_classic()
内容的提问来源于stack exchange,提问作者Shawn Hemelstrand
相关产品推荐
相关产品推荐

