在R中实现逐行与指定行的批量Logistic回归分析(循环方案问询)
Absolutely! Ditching those repetitive manual glm calls for a loop (or vectorized function) is exactly the right move when you're dealing with 6000 rows—no more copy-pasting madness. Let's walk through how to do this efficiently.
First off, we need to adjust your data setup. Right now, your r1-r6 and C1 are row names in dat, but logistic regression expects dependent/independent variables to be columns in your data frame. So we'll transpose the data to flip rows and columns, turning those row names into usable column names:
# Transpose the data frame and convert back to a data frame (preserves row/col names) dat_t <- as.data.frame(t(dat)) # Check the new structure - now r1-r6 and C1 are columns, X1-X30 are rows (our observations) head(dat_t)
lapply (Clean, R-Style Approach) Using lapply is the most concise way to handle this. We'll target all the response columns (r1-r6, or all columns except C1 in your full dataset), fit a logistic regression for each, and extract the standard errors:
# Get all column names except C1 (these are our dependent variables) response_cols <- setdiff(colnames(dat_t), "C1") # Batch-fit models and extract standard errors se_results <- lapply(response_cols, function(col) { # Build the formula dynamically: [dependent variable] ~ C1 model_formula <- as.formula(paste(col, "~ C1")) # Fit logistic regression glm_model <- glm(model_formula, data = dat_t, family = binomial) # Pull standard errors for both intercept and C1 coefficient coef(summary(glm_model))[, 2] }) # Convert the list of results into a tidy data frame se_df <- as.data.frame(do.call(rbind, se_results)) rownames(se_df) <- response_cols colnames(se_df) <- c("Intercept_SE", "C1_SE")
If you prefer a more explicit approach, a for loop works just as well:
# Initialize an empty data frame to store results se_df <- data.frame( Intercept_SE = numeric(), C1_SE = numeric(), row.names = character() ) # Get list of dependent variable columns response_cols <- setdiff(colnames(dat_t), "C1") # Loop through each dependent variable for (col in response_cols) { model_formula <- as.formula(paste(col, "~ C1")) glm_model <- glm(model_formula, data = dat_t, family = binomial) # Extract standard errors and add to results data frame se_vals <- coef(summary(glm_model))[, 2] se_df[col, ] <- se_vals }
With 6000 models, you might run into edge cases (like complete separation in the data that breaks glm). Add tryCatch to skip errors and return NA instead of crashing the whole process:
se_results <- lapply(response_cols, function(col) { tryCatch({ model_formula <- as.formula(paste(col, "~ C1")) glm_model <- glm(model_formula, data = dat_t, family = binomial) coef(summary(glm_model))[, 2] }, error = function(e) { # Return NA if model fails to fit c(NA, NA) }) })
内容的提问来源于stack exchange,提问作者SUMIT

