You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R的geepack包中处理TB纵向数据缺失值的GEE模型报错问题

Fixing GEE Missing Value Errors in geepack (R) for TB Longitudinal Data

Got it, let's work through this missing value issue with your TB patient longitudinal dataset and the geepack package—this is a super common pain point with repeated measures data, so I’ve got actionable steps to get your GEE model running smoothly.

First, let’s recap why this happens: By default, geeglm() uses listwise deletion (na.action = na.omit), which drops any entire visit row if any predictor has a missing value. For TB datasets, where some tests might be skipped at certain visits but others are fully recorded, this can unnecessarily reduce your sample size or even break the model if too many rows are dropped.

Here’s how to fix it:

1. First, Map Your Missing Data Pattern

Before jumping into fixes, you need to understand where and how much data is missing. This helps you pick the right strategy:

# Quick summary of missing values per variable
summary(your_tb_dataset)

# Visualize missing patterns (install mice if you haven't)
library(mice)
md.pattern(your_tb_dataset)

Look for variables with high missing rates, or if missingness is clustered in specific visits/patients. This tells you if missingness is random (MAR/MCAR) or systematic (MNAR)—which guides your next move.

2. Fix Non-Standard Missing Values

Sometimes missing data is stored as strings like "NA", "Missing", or blank spaces instead of true NA values. geepack will treat these as valid factor levels, causing weird errors. Fix this first:

# Using dplyr (install if needed) to convert string missing to NA
library(dplyr)
your_tb_dataset <- your_tb_dataset %>%
  mutate(across(everything(), ~ifelse(. %in% c("NA", "Missing", ""), NA, .)))

# Convert numeric variables back if they got coerced to character
your_tb_dataset$age <- as.numeric(your_tb_dataset$age)

3. Choose a Missing Value Handling Strategy

Option A: Keep More Rows with na.exclude

If you want to avoid dropping entire visits unless absolutely necessary, switch the na.action parameter to na.exclude. This retains the structure of your dataset but ignores missing rows during model calculation:

library(geepack)
gee_model <- geeglm(
  outcome ~ predictor1 + predictor2 + visit_time,
  data = your_tb_dataset,
  id = patient_id,  # Your clustering variable (patient ID)
  family = binomial,  # Binary outcome: improved/not improved
  corstr = "exchangeable",  # Pick the right correlation structure for your data
  na.action = na.exclude
)

Note: This still drops visits with missing values, but preserves the original dataset row count in output (useful for diagnostics).

Option B: Impute Missing Values (Best for Retaining Data)

For longitudinal data, multiple imputation is the gold standard—especially if missingness is random. Use the mice package to impute values while accounting for the repeated measures structure:

# Run multiple imputation (adjust method based on variable type)
imputed_datasets <- mice(
  your_tb_dataset,
  method = "pmm",  # Predictive mean matching for continuous variables; use "logreg" for binary
  maxit = 5,  # Number of imputation iterations
  seed = 123,  # For reproducibility
  m = 5  # Number of imputed datasets
)

# Run GEE on each imputed dataset and pool results
library(mitml)
imputed_list <- mitmlComplete(imputed_datasets, "all")

# Define a function to run GEE on each dataset
run_gee <- function(data) {
  geeglm(
    outcome ~ predictor1 + predictor2 + visit_time,
    data = data,
    id = patient_id,
    family = binomial,
    corstr = "exchangeable"
  )
}

# Apply function and pool results
gee_results <- lapply(imputed_list, run_gee)
pooled_results <- testEstimates(gee_results, var.comp = "wald")

# View final pooled model
summary(pooled_results)

This method uses all your data and accounts for uncertainty from imputation.

Option C: Treat Missingness as a Separate Category (For Categorical Predictors)

If a categorical predictor has missing values with a clear clinical meaning (e.g., "no test performed"), you can recode missingness as its own factor level instead of dropping the row:

# Convert missing values to a new factor level
your_tb_dataset$test_result <- as.factor(
  ifelse(is.na(your_tb_dataset$test_result), "Not Tested", as.character(your_tb_dataset$test_result))
)

# Now run GEE as usual
gee_model <- geeglm(
  outcome ~ predictor1 + test_result + visit_time,
  data = your_tb_dataset,
  id = patient_id,
  family = binomial,
  corstr = "exchangeable"
)

Only use this if the missing category has a meaningful interpretation—don’t do this for continuous variables.

Final Checks

After fixing missing values, double-check that your model runs without errors and that the sample size makes sense. If you still get errors, verify that your clustering variable (id) is correctly formatted as a factor or numeric identifier—sometimes weird formatting here can cause issues alongside missing data.

内容的提问来源于stack exchange,提问作者Dilsher Singh Dhillon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 08:32:57