如何修正glmmTMB零膨胀过分散数据的DHARMa分位数偏差
鸟类降落数据建模问题梳理
数据背景
- 分析目标:探究特定特征鸟类是否更倾向于在特定区域降落
- 数据集:19447条单日观测汇总记录(截至2022年),覆盖12个池塘,因此
size_ha(池塘面积)和River_prox(距河距离)存在重复值 - 数据特性:严重零膨胀+过分散,其中一个池塘的鸟类降落量约为其他池塘的10倍;年检测频率上升进一步加剧了分散问题
模型尝试
采用glmmTMB包建模,分别尝试poisson和negbin2族模型,认为负二项模型更适配过分散数据。模型纳入池塘面积的偏移项offset(log(Size_ha)),并尝试将Pond_ID、Year、DOY(日序)作为随机效应,但不确定最优结构设置;考虑到池塘面积逐年变化,认为Pond_ID应嵌套于Year,但对此设置存疑。
Poisson零膨胀模型
model_full <- glmmTMB(Landed ~ Guild * River_prox + offset(log(Size_ha))+(1|DOY) + (1|Year/Pond_ID), ziformula= ~Guild,data=all_landings_sum_guild, family="poisson")
模型摘要:
AIC BIC logLik deviance df.resid 295482.5 295646.5 -147717.2 295434.5 6855
负二项模型(调整随机效应结构版)
因原随机效应结构报错,调整后模型代码:
model <- glmmTMB(formula = Landed ~ offset(log(Size_ha)) + Guild * River_prox + (1 + (1 | Year/Pond_ID) + (1 | DOY)), family = nbinom2, data = all_landings_sum_guild)
模型摘要:
AIC BIC logLik deviance df.resid 50610.9 50734.0 -25287.5 50574.9 6861
模型拟合问题
随机效应结构报错
使用(1|DOY)+(1|Year/Pond_ID)结构时出现以下警告:
Warning messages:
1: In finalizeTMB(TMBStruc, obj, fit, h, data.tmb.old) :
Model convergence problem; non-positive-definite Hessian matrix. See vignette('troubleshooting')
2: In finalizeTMB(TMBStruc, obj, fit, h, data.tmb.old) :
Model convergence problem; function evaluation limit reached without convergence (9). See vignette('troubleshooting'), help('diagnose')
已查阅故障排除文档,调整效应结构可消除警告,但残差图表现无明显改善。
DHARMa拟合诊断结果(基于负二项模型)
诊断代码:
simulationOutput <- simulateResiduals(fittedModel = model, plot = FALSE) plot(simulationOutput) testQuantiles(simulationOutput) testZeroInflation(simulationOutput) testOutliers(simulationOutput) testDispersion(simulationOutput)
各测试结果:
- 分位数检验:
data: simulationOutput p-value < 2.2e-16 alternative hypothesis: both
- 零膨胀检验:
data: simulationOutput ratioObsSim = 0.042636, p-value < 2.2e-16 alternative hypothesis: two.sided
- 异常值检验:
outliers at both margin(s) = 72, observations = 6879, p-value = 0.02487 alternative hypothesis: true probability of success is not equal to 0.007968127 95 percent confidence interval: 0.008198258 0.013163080 sample estimates: frequency of outliers (expected: 0.00796812749003984 ) 0.01046664
- 分散性检验:
data: simulationOutput dispersion = 1.294, p-value = 0.4 alternative hypothesis: two.sided
- Bootstrap异常值检验:
outliers at both margin(s) = 71, observations = 6879, p-value = 0.22
备择假设:双侧
疑问与困境
- 不确定是否需要全部纳入
Pond_ID、Year、DOY作为随机效应,以及如何设置最优结构 - 因池塘面积逐年变化,认为
Pond_ID嵌套于Year,但不确定该如何在模型中正确设置该结构 - 已查阅
DHARMa、glmmTMB故障排除指南及论坛方案,但要么无法理解操作,要么执行报错,问题未解决;怀疑对数据特性存在认知盲区
内容的提问来源于stack exchange,提问作者Hanna Ulmer
相关产品推荐
相关产品推荐

