循环运行多模型时遇invalid type (list)错误:如何指定因变量/自变量?
解决polr模型动态指定变量的报错问题
问题背景
需要创建bootstrapped数据集,生成IPW权重后,通过嵌套循环自动化运行80余个不同因变量/自变量的polr模型,但建模部分持续报错。
错误信息
Error in model.frame.default(formula = boot.samples[[i]][, outcome] ~ :
invalid type (list) for variable 'boot.samples[[i]][, predictors]'
原代码问题点
当predictors是包含多个变量名的向量时,boot.samples[[i]][, predictors]会返回一个数据框(list类型),而polr的公式参数无法直接接受list类型的变量输入,必须用标准的公式表达式(如Severity ~ Weight + Age)。
解决方案
核心是动态构建公式字符串,再转换为公式对象传入polr,同时优化模型结果的存储逻辑(避免覆盖之前bootstrap样本的结果)。
修改后的完整代码
# Creation of mock data library(MASS) # 需加载polr所在的MASS包 data <- data.frame( "Severity" = as.factor(c(rep("None", 25), rep("Mild", 25), rep("Moderate", 25), rep("Severe", 25))), "Severity2" = as.factor(c(rep("None", 40), rep("Mild", 20), rep("Moderate", 20), rep("Severe", 20))), "Weight" = rnorm(100, mean = 160, sd = 30), "Age" = rnorm(100, mean = 40, sd = 7), "Gender" = as.factor(rbinom(100, size = 1, prob = 0.5)), "Tested" = as.factor(rbinom(100, size = 1, prob = 0.4)) ) data$Severity <- ifelse(data$Tested == 0, NA, data$Severity) data$Severity2 <- ifelse(data$Tested == 0, NA, data$Severity2) data$Severity <- ordered(data$Severity, levels = c("None", "Mild", "Moderate", "Severe")) data$Severity2 <- ordered(data$Severity2, levels = c("None", "Mild", "Moderate", "Severe")) # Creating boostrapped datasets nboot <- 2 set.seed(10) boot.samples <- lapply(1:nboot, function(i) { data[base::sample(1:nrow(data), replace = TRUE),] }) # Create empty list to store results later - 改为嵌套列表,按bootstrap样本+模型维度存储 coefs <- vector("list", length(boot.samples)) for (i in seq_along(coefs)) { coefs[[i]] <- vector("list", length(models)) } # Setting up the outcomes/predictors of each of the models I will run # 优化模型定义:用命名列表更清晰 mod1 <- list(outcome = "Severity", preds = c("Weight","Age")) mod2 <- list(outcome = "Severity2", preds = c("Weight", "Age", "Gender")) models <- list(mod1, mod2) # Running the for-loop for(i in 1:length(boot.samples)) { #Setting up weight creation null <- glm(formula = Tested ~ 1, family = "binomial", data = boot.samples[[i]]) full <- glm(formula = Tested ~ Age, family = "binomial", data = boot.samples[[i]]) step <- step(null, k = 2, direction = "forward", scope=list(lower = null, upper = full), trace = 0) pd.combined <- stats::predict(step, type = "response") numer.combined <- glm(Tested ~ 1, family = "binomial", data = boot.samples[[i]]) pn.combined <- stats::predict(numer.combined, type = "response") # Creating stabilized weights boot.samples[[i]]$ipw <- ifelse(boot.samples[[i]]$Tested==0, ((1-pn.combined)/(1-pd.combined)), (pn.combined)/(pd.combined)) # Now running each model and storing the coefficients for(j in 1:length(models)) { outcome <- models[[j]]$outcome # 用命名列表提取更直观 predictors <- models[[j]]$preds # 动态构建公式字符串:outcome ~ pred1 + pred2 + ... formula_str <- paste(outcome, paste(predictors, collapse = " + "), sep = " ~ ") formula_obj <- as.formula(formula_str) # 直接用data参数传入bootstrap样本,避免重复索引 model_results <- polr(formula = formula_obj, data = boot.samples[[i]], weights = ipw, # 直接用变量名,data参数会自动匹配 method = "logistic", Hess = TRUE) # 按bootstrap样本+模型的维度存储系数 coefs[[i]][[j]] <- model_results$coefficients } }
关键修改点
- 动态构建公式:
用paste()将因变量和自变量拼接成公式字符串(如"Severity ~ Weight + Age"),再通过as.formula()转换为公式对象,这是polr可识别的标准格式。 - 优化数据传入方式:
给polr传入data = boot.samples[[i]],直接在公式中使用变量名,避免手动索引数据框,代码更简洁且不易出错。 - 结果存储优化:
将coefs改为嵌套列表,每个bootstrap样本对应一个子列表,存储该样本下所有模型的系数,避免原代码中后续bootstrap样本覆盖前序样本结果的问题。 - 模型定义优化:
用命名列表(outcome = "Severity")替代匿名列表,代码可读性更强。
内容的提问来源于stack exchange,提问作者kjb17
相关产品推荐
相关产品推荐

