多重插补嵌套数据集分组后执行线性回归的问题求助
问题分析与解决方案
核心错误点
代码2(map方式)的问题
你在map(lm, formula = wgt ~ bmi)这一步没有实现按reg分组执行回归:group_by(reg)只是给数据集添加了分组标记,但lm()不会自动遍历分组,最终每个插补数据集只跑了一个全局回归,而非分reg的分组回归,所以pool()后只能得到整体的模型结果,丢失了reg分组信息。
代码3(for循环方式)的问题
do(lm(...))返回的是lm模型对象,但do()要求每个分组必须返回数据框,因此触发类型错误;data = boys_imp2[[i]]错误地引用了整个插补数据集,而非当前分组的子数据,应该用.指代do()中的当前分组数据。
正确实现方法
我们需要对每个插补数据集按reg分组跑回归,整理结果后再合并多重插补的估计值,最终得到和非插补数据格式一致的分组模型结果:
# 加载依赖包 library(mice) library(tidyverse) # 生成多重插补数据集 boys_imp <- mice(boys, printFlag = FALSE) # 步骤1:遍历每个插补数据集,按reg分组跑回归并整理成tidy格式 imputed_tidy_results <- boys_imp %>% mice::complete("all") %>% map(function(imp_df) { imp_df %>% group_by(reg) %>% do(tidy(lm(wgt ~ bmi, data = .), conf.int = TRUE)) }) # 步骤2:合并所有插补结果,添加插补编号标记 combined_results <- bind_rows(imputed_tidy_results, .id = "imputation") # 步骤3:按reg和term分组,合并多重插补的估计值 pooled_group_results <- combined_results %>% group_by(reg, term) %>% summarise( # 用pool.scalar合并多重插补的估计值与标准误 estimate = pool.scalar(estimate, std.error)$qbar, std.error = pool.scalar(estimate, std.error)$t, df = pool.scalar(estimate, std.error)$df, # 计算统计量、p值和置信区间 statistic = estimate / std.error, p.value = 2 * pt(-abs(statistic), df = df), conf.low = estimate - qt(0.975, df) * std.error, conf.high = estimate + qt(0.975, df) * std.error, .groups = "drop" ) # 查看最终结果 print(pooled_group_results, n = Inf)
结果说明
最终输出的pooled_group_results格式和你期望的summary(mod1)完全一致,包含每个reg分组下的截距项、bmi系数的估计值、标准误、统计量、p值及置信区间,同时已经完成了多重插补结果的合并(符合Rubin规则)。
内容的提问来源于stack exchange,提问作者Mohamed Yusuf
相关产品推荐
相关产品推荐

