You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

循环运行多模型时遇invalid type (list)错误:如何指定因变量/自变量?

解决polr模型动态指定变量的报错问题

问题背景

需要创建bootstrapped数据集,生成IPW权重后,通过嵌套循环自动化运行80余个不同因变量/自变量的polr模型,但建模部分持续报错。

错误信息

Error in model.frame.default(formula = boot.samples[[i]][, outcome] ~ :
invalid type (list) for variable 'boot.samples[[i]][, predictors]'

原代码问题点

当predictors是包含多个变量名的向量时,boot.samples[[i]][, predictors]会返回一个数据框(list类型),而polr的公式参数无法直接接受list类型的变量输入,必须用标准的公式表达式(如Severity ~ Weight + Age)。


解决方案

核心是动态构建公式字符串,再转换为公式对象传入polr,同时优化模型结果的存储逻辑(避免覆盖之前bootstrap样本的结果)。

修改后的完整代码

# Creation of mock data
library(MASS) # 需加载polr所在的MASS包

data <- data.frame(
  "Severity" = as.factor(c(rep("None", 25), rep("Mild", 25), rep("Moderate", 25), rep("Severe", 25))),
  "Severity2" = as.factor(c(rep("None", 40), rep("Mild", 20), rep("Moderate", 20), rep("Severe", 20))),
  "Weight" = rnorm(100,  mean = 160, sd = 30),
  "Age" = rnorm(100,  mean = 40, sd = 7),
  "Gender" = as.factor(rbinom(100, size = 1, prob = 0.5)),
  "Tested" = as.factor(rbinom(100, size = 1, prob = 0.4))
)

data$Severity <- ifelse(data$Tested == 0, NA, data$Severity)
data$Severity2 <- ifelse(data$Tested == 0, NA, data$Severity2)

data$Severity <- ordered(data$Severity, levels = c("None", "Mild", "Moderate", "Severe"))
data$Severity2 <- ordered(data$Severity2, levels = c("None", "Mild", "Moderate", "Severe"))

# Creating boostrapped datasets
nboot <- 2
set.seed(10)
boot.samples <- lapply(1:nboot, function(i) {
  data[base::sample(1:nrow(data), replace = TRUE),]
})

# Create empty list to store results later - 改为嵌套列表,按bootstrap样本+模型维度存储
coefs <- vector("list", length(boot.samples))
for (i in seq_along(coefs)) {
  coefs[[i]] <- vector("list", length(models))
}

# Setting up the outcomes/predictors of each of the models I will run
# 优化模型定义:用命名列表更清晰
mod1 <- list(outcome = "Severity", preds = c("Weight","Age"))
mod2 <- list(outcome = "Severity2", preds = c("Weight", "Age", "Gender"))
models <- list(mod1, mod2)

# Running the for-loop
for(i in 1:length(boot.samples)) {
  
  #Setting up weight creation
  null <- glm(formula = Tested ~ 1, family = "binomial", data = boot.samples[[i]])
  full <- glm(formula = Tested ~ Age, family = "binomial", data = boot.samples[[i]])
  step <- step(null, k = 2, direction = "forward", scope=list(lower = null, upper = full),  trace = 0)
  pd.combined <- stats::predict(step, type = "response")
  numer.combined <- glm(Tested ~ 1, family = "binomial", 
                        data = boot.samples[[i]])
  pn.combined <- stats::predict(numer.combined, type = "response") 
  
  # Creating stabilized weights
  boot.samples[[i]]$ipw <- ifelse(boot.samples[[i]]$Tested==0, ((1-pn.combined)/(1-pd.combined)), (pn.combined)/(pd.combined))
  
  # Now running each model and storing the coefficients
  for(j in 1:length(models)) {   
    outcome <- models[[j]]$outcome # 用命名列表提取更直观
    predictors <- models[[j]]$preds
    
    # 动态构建公式字符串:outcome ~ pred1 + pred2 + ...
    formula_str <- paste(outcome, paste(predictors, collapse = " + "), sep = " ~ ")
    formula_obj <- as.formula(formula_str)
    
    # 直接用data参数传入bootstrap样本,避免重复索引
    model_results <- polr(formula = formula_obj, 
                          data = boot.samples[[i]], 
                          weights = ipw, # 直接用变量名,data参数会自动匹配
                          method = "logistic", 
                          Hess = TRUE)
    
    # 按bootstrap样本+模型的维度存储系数
    coefs[[i]][[j]] <- model_results$coefficients
  }
}

关键修改点

  • 动态构建公式:
    用paste()将因变量和自变量拼接成公式字符串(如"Severity ~ Weight + Age"),再通过as.formula()转换为公式对象,这是polr可识别的标准格式。
  • 优化数据传入方式:
    给polr传入data = boot.samples[[i]],直接在公式中使用变量名,避免手动索引数据框,代码更简洁且不易出错。
  • 结果存储优化:
    将coefs改为嵌套列表,每个bootstrap样本对应一个子列表,存储该样本下所有模型的系数,避免原代码中后续bootstrap样本覆盖前序样本结果的问题。
  • 模型定义优化:
    用命名列表(outcome = "Severity")替代匿名列表,代码可读性更强。

内容的提问来源于stack exchange,提问作者kjb17

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 20:20:52