如何基于mice包插补数据后获取lm模型的标准化Beta系数?
Great question! When working with multiply imputed data from the mice package, getting standardized beta weights does require a bit of extra work since the pool() function only returns unstandardized coefficients by default. Here are two reliable approaches to solve this:
Method 1: Standardize Continuous Variables Before Imputation
The most straightforward way is to standardize your continuous predictor and outcome variables (mean = 0, standard deviation = 1) before running the imputation. This way, when you fit your linear model on the imputed standardized data, the pooled coefficients will directly be the standardized beta weights.
Step-by-Step Code:
# Load required package library(mice) # Example dataset (replace with your data) set.seed(123) dat <- data.frame( y = rnorm(100), x1 = rnorm(100), x2 = rnorm(100), x3 = sample(c(1,2,3), 100, replace = TRUE) # Categorical variable (no standardization needed) ) # Add missing values dat$y[sample(1:100, 10)] <- NA dat$x1[sample(1:100, 15)] <- NA # Standardize only continuous variables cont_vars <- sapply(dat, is.numeric) dat_std <- dat dat_std[, cont_vars] <- scale(dat_std[, cont_vars]) # Run multiple imputation on standardized data imp_std <- mice(dat_std, m = 5, seed = 123) # Fit linear model on each imputed dataset fit_std <- with(imp_std, lm(y ~ x1 + x2 + x3)) # Pool results and view standardized betas pooled_std <- pool(fit_std) summary(pooled_std)
The coefficients in the output are now standardized betas. Note that categorical variables (like x3 in the example) don't get standardized, but their coefficients represent the standardized effect relative to the reference category, which is still interpretable.
Method 2: Calculate Standardized Betas Manually From Pooled Results
If you already have a pooled model from unstandardized data, you can compute standardized betas using the formula:
Standardized Beta = (Unstandardized Beta * SD of Predictor) / SD of Outcome
Step-by-Step Code:
# First, run imputation and model on unstandardized data (if not already done) imp <- mice(dat, m = 5, seed = 123) fit <- with(imp, lm(y ~ x1 + x2 + x3)) pooled <- pool(fit) # Extract unstandardized coefficients coef_unstd <- coef(pooled) # Calculate standard deviations using the pooled imputed data (more accurate than raw data with NAs) completed_dat <- complete(imp, "long") # Combine all imputed datasets sd_x_pooled <- sapply(c("x1", "x2"), function(var) { mean(tapply(completed_dat[[var]], completed_dat$.imp, sd)) # Average SD across imputations }) sd_y_pooled <- mean(tapply(completed_dat$y, completed_dat$.imp, sd)) # Compute standardized betas (exclude intercept) coef_std <- coef_unstd[-1] coef_std[c("x1", "x2")] <- coef_std[c("x1", "x2")] * sd_x_pooled / sd_y_pooled coef_std <- c(0, coef_std) # Intercept is 0 in standardized models names(coef_std) <- names(coef_unstd) # View results coef_std
Key Notes:
- Using the pooled imputed data to calculate SDs accounts for uncertainty from missing values, which is better than using raw data with
na.rm = TRUE. - This method gives you the standardized betas, but you won't get automatically calculated standard errors or p-values for them. For that, Method 1 is better.
内容的提问来源于stack exchange,提问作者Zach H

