如何将逻辑回归模型的OR、CI、p值及协变量信息导出为数据框并保存至Excel
问题:Logistic回归模型输出丢失变量/水平标签的解决方法
我正在构建多个logistic regression模型,需要导出OR(优势比)、CI(置信区间)、p值以及协变量/水平信息。目前已成功将OR、CI、p值导出为数据框,但变量/水平的标签在导出过程中丢失。
原R代码如下:
#packages library(tidyverse) install.packages("AER") library("AER") library(writexl) #data data(Affairs, package="AER") Affairs$ynaffair[Affairs$affairs > 0] <- 1 Affairs$ynaffair[Affairs$affairs == 0] <- 0 # logistic regression model model <- glm(ynaffair~gender + age + yearsmarried + children + religiousness + education + occupation + rating, family = binomial, data = Affairs) summary(model) #formatting the output model_output <- as.data.frame(cbind(round(exp(model$coefficients), 2), exp(confint.default(model)), summary(model)$coefficients[,4])) %>% mutate_if(is.numeric, round, digits = 3) %>% unite(CI, c(`2.5 %`, `97.5 %`), sep = ", ", remove = TRUE) # Exporting it to Excel write_xlsx(model_output, "model_output.xlsx")
解决方案
原代码中,变量/水平标签其实是以行名的形式存在于model_output数据框中,但write_xlsx默认不会导出行名,导致标签丢失。我们只需要把行名转换为数据框的正式列,就能保留标签信息。
修改后的完整代码:
#packages library(tidyverse) install.packages("AER") library("AER") library(writexl) #data data(Affairs, package="AER") Affairs$ynaffair <- ifelse(Affairs$affairs > 0, 1, 0) # 简化赋值逻辑 # logistic regression model model <- glm(ynaffair~gender + age + yearsmarried + children + religiousness + education + occupation + rating, family = binomial, data = Affairs) #formatting the output model_output <- tibble( Covariate = names(model$coefficients), # 提取变量/水平标签作为单独列 OR = round(exp(model$coefficients), 3), CI_low = exp(confint.default(model))[,1], CI_high = exp(confint.default(model))[,2], p_value = summary(model)$coefficients[,4] ) %>% mutate( CI_low = round(CI_low, 3), CI_high = round(CI_high, 3), p_value = round(p_value, 3), CI = str_c(CI_low, CI_high, sep = ", ") # 合并置信区间 ) %>% select(Covariate, OR, CI, p_value) # 调整列顺序 # Exporting it to Excel write_xlsx(model_output, "model_output_with_labels.xlsx")
关键修改点
- 用
tibble直接构建数据框,显式添加Covariate列存储变量/水平标签,避免依赖行名(行名不会被write_xlsx自动导出) - 简化
ynaffair的赋值逻辑,用ifelse替代多次索引赋值,代码更简洁 - 分步计算OR、置信区间上下限和p值,再合并CI,逻辑更清晰易维护
- 最后通过
select调整列顺序,确保输出结构符合需求
内容的提问来源于stack exchange,提问作者Newtostats_24
相关产品推荐
相关产品推荐

