gtsummary::tbl_regression返回异常:硕士学位行系数格式错误
问题描述
用户使用gtsummary::tbl_regression格式化逻辑回归结果时出现异常:
dat.mi <- import(here("temp", "adTurn_mi.rds")) fit1 <- glm(quit ~ sex + age + race + education + experience + tenure_sr + salary + anymc + catcap + medi + nonpro + rural, family = 'binomial', weights = svy_weight, data = complete(dat.mi) ) |> tbl_regression(exponentiate = TRUE)
dat.mi是经complete()补全的MICE数据- 多数变量的OR值(指数化系数)显示正常,但“硕士学位”行输出逗号分隔的数字串,而非正确的OR值
解决方法
- 检查变量编码
确认education变量中“硕士学位”对应的水平是否为因子类型且编码正确。若为字符型变量,先转为因子并指定合理水平顺序:
dat.mi$education <- factor(dat.mi$education, levels = c("低学历", "本科", "硕士学位")) # 替换为实际水平
- 强制指定变量处理逻辑
在tbl_regression()中针对education变量单独设置,确保系数被正确指数化:
fit1 <- glm(...) |> tbl_regression( exponentiate = TRUE, modify = list( update_table_body( mutate, estimate = ifelse(variable == "education" & label == "硕士学位", exp(estimate), estimate), conf.low = ifelse(variable == "education" & label == "硕士学位", exp(conf.low), conf.low), conf.high = ifelse(variable == "education" & label == "硕士学位", exp(conf.high), conf.high) ) ) )
- 修正表格异常值
若上述方法无效,可直接修改表格的table_body部分,替换异常的逗号分隔值:
fit1$table_body <- fit1$table_body |> mutate( across(c(estimate, conf.low, conf.high), ~ifelse(label == "硕士学位", exp(as.numeric(str_replace_all(., ",", ""))), .)) )
- 确认参考水平设置
如果education是多分类变量,检查是否正确设置了参考水平,避免系数估计异常:
dat.mi$education <- relevel(dat.mi$education, ref = "本科") # 设置合理参考水平
内容的提问来源于stack exchange,提问作者Nate P
相关产品推荐
相关产品推荐

