编写多元二项Logistic回归分析循环及对应数据框技术咨询
Got it, let's tackle this problem step by step. You want to run multiple binary logistic regression models in a loop using your data frame, where disease is the binary outcome variable. Here's how you can do it properly in R:
Binary Logistic Regression Loop for Your Data Frame
First, let's confirm your data frame is set up correctly (I'll include it here for reference):
df <- data.frame( disease = c(0,1,0,1), var1 = c(0,1,2,0), var2 =c(0,1,2,0), var3 = c(0,1,2,0), var40 = c(0,1,2,0), Bi = c(0,1,0,1), gender = c(1,0,1,0), P1 = c(-0.040304832,0.006868288,0.002663759,0.020251087), P2 = c(0.010566526,0.002663759,0.017480721,-0.008685749), P3 = c(-0.008685749,0.020251087,-0.040304832,0.002663759), P4 = c(0.017480721,0.024306667,0.002663759,0.010566526), stringsAsFactors = FALSE )
Option 1: Loop Through Single Predictors + Control Variables
If you want to run a model for each predictor while including a fixed set of control variables (e.g., Bi and gender), this code will work perfectly:
# Define control variables and list of predictors to iterate over control_vars <- c("Bi", "gender") predictors <- setdiff(names(df), c("disease", control_vars)) # Initialize a list to store all model results (clean way to keep track) model_list <- list() # Loop through each predictor for (pred in predictors) { # Build the model formula dynamically formula_str <- paste("disease ~", pred, "+", paste(control_vars, collapse = " + ")) formula <- as.formula(formula_str) # Fit the binary logistic regression model model <- glm(formula, data = df, family = binomial(link = "logit")) # Store the model in the list with a descriptive name model_list[[paste0("model_", pred)]] <- model # Optional: Print summary for each model as it runs (great for quick checks) cat("\n--- Summary for Model with Predictor:", pred, "---\n") print(summary(model)) } # Later, you can access individual models like this: model_list$model_var1
Option 2: Loop Through Multivariate Model Variations
If you want to test different full multivariate models (e.g., exclude one predictor at a time from the full model), here's how to do that:
# Get all predictor variables (excluding the outcome) all_predictors <- setdiff(names(df), "disease") # Initialize list to store these models full_model_list <- list() # First, fit the full model with all predictors full_formula <- as.formula(paste("disease ~", paste(all_predictors, collapse = " + "))) full_model_list[["full_model"]] <- glm(full_formula, data = df, family = binomial(link = "logit")) cat("\n--- Summary for Full Multivariate Model ---\n") print(summary(full_model_list[["full_model"]])) # Now loop to fit models excluding one predictor at a time for (pred in all_predictors) { # Create a subset of predictors excluding the current one subset_predictors <- setdiff(all_predictors, pred) formula_str <- paste("disease ~", paste(subset_predictors, collapse = " + ")) formula <- as.formula(formula_str) # Fit the model model <- glm(formula, data = df, family = binomial(link = "logit")) full_model_list[[paste0("model_excluding_", pred)]] <- model # Print summary cat("\n--- Summary for Model Excluding:", pred, "---\n") print(summary(model)) }
Quick Notes to Keep in Mind
- Categorical Variables: Your categorical variables like
Biandgenderare already coded as 0/1, which works forglm. If you had unordered factors with more levels, convert them to factors first withdf$var <- as.factor(df$var)to generate proper dummy variables. - Sample Size: Heads up—your dataset only has 4 rows, which is way too small for reliable logistic regression results. This code is for demonstration; you'll need much more data in practice to get meaningful coefficients and p-values.
- Result Storage: Using lists to store models is efficient—you can easily extract coefficients, p-values, or run further analyses on all models at once (e.g.,
lapply(model_list, coef)to get all coefficients in one go).
内容的提问来源于stack exchange,提问作者Maneesh Mishra
相关产品推荐
相关产品推荐

