使用gamlss建模freq时遇while循环缺失值错误,求解决
Let's break down your problem and fix that error step by step:
Your Scenario
You're working with a dataset where the response variable freq consists of proportional values (0.48, 0.74, etc. — all between 0 and 1). You tried fitting GAMLSS models using Poisson (PO()) and Negative Binomial (NBI()) distributions, but hit this error:
Error in while (abs(G.dev.old - G.dev) > c.crit && iter < n.cyc) { : missing value where TRUE/FALSE needed
Root Cause
The key issue here is mismatched distribution assumptions:
- Poisson and Negative Binomial distributions are designed for count data (non-negative integers like the number of events).
- Your
freqvariable is a continuous proportion (0 < freq < 1), which violates the count data requirement. This mismatch causes numerical instability during model fitting, leading to NA values in the deviance calculation — hence thewhileloop error (it can't evaluate TRUE/FALSE when values are missing).
Fixes to Try
1. Use a Distribution Designed for Proportional Data
The most appropriate choice here is the Beta distribution (available in GAMLSS as BE()), which is built specifically for 0-1 proportional responses. Adjust your code like this:
# Fit Beta GAMLSS model with logit link (standard for proportions) model_beta <- gamlss(freq ~ Sexo + zona_2016 + rango_edad, family = BE(mu.link = "logit"), # Logit link maps linear predictor to 0-1 range data = na.omit(subset(datos, !is.na(freq))))
- The logit link is the default for
BE(), but you can also useprobitorcloglogif your data warrants it. - If your
freqincludes exact 0s or 1s, consider using a zero/one-inflated Beta distribution (BEZI()) instead.
2. Reorient the Model to Use Count Data (If Appropriate)
Looking at your dataset, you have siniestros (count of incidents) and expuestos (exposure count). If freq is calculated as siniestros / expuestos, you should model the count of incidents directly with an offset for exposure — this aligns with Poisson/Negative Binomial use cases:
# Model incident counts with exposure as an offset model_nbi <- gamlss(siniestros ~ Sexo + zona_2016 + rango_edad + offset(log(expuestos)), family = NBI(), data = na.omit(subset(datos, !is.na(siniestros) & !is.na(expuestos))))
This approach properly accounts for the exposure size and uses the count distribution as intended.
3. Debug with Simplified Models
If you still run into issues, narrow down the problem by starting with a simple model:
# Test a single predictor first simple_model <- gamlss(freq ~ Sexo, family = BE(), data = na.omit(subset(datos, !is.na(freq))))
If this works, gradually add variables (zona_2016, then rango_edad) to identify if a specific variable is causing instability (e.g., rare categories with sparse data).
4. Validate Data Quality
Double-check your dataset for:
- Out-of-range values in
freq(values outside 0-1 would break Beta distribution fitting) - Sparse categories in predictors (e.g., very small counts in
zona_2016= "Alejada") which can cause numerical issues - Missing values that might have slipped through your
na.omit()call
内容的提问来源于stack exchange,提问作者bubleskmy

