R中glm自定义单变量逻辑回归函数无法运行问题求助
Fixing Your Univariate Logistic Regression Code
Let's walk through the issues in your current code and get it working smoothly.
Key Problems in Your Original Code
- Incorrect
lapplysyntax: You’re assigning the functionunivariate_logisticinside thelapplycall, which isn’t the standard or reliable way to pass a function tolapply. While R might try to evaluate this, it’s prone to unexpected behavior. - Manual variable list: Typing
A1toA50by hand is error-prone and unnecessary—we can generate this list programmatically in one line. - Not storing results: If you don’t assign the output of
lapplyto a variable, you won’t be able to access or review the regression summaries later.
Corrected Code
First, generate your predictor list cleanly to avoid typos:
# Create list of predictor variables (A1 to A50) predictors <- paste0("A", 1:50)
Then, use one of these two approaches to run your regressions:
Option 1: Anonymous Function Directly in lapply
This is concise for one-off use:
# Run regressions and store all summaries univariate_results <- lapply(predictors, function(var) { # Build the formula dynamically formula <- as.formula(paste("churn_flag ~", var)) # Fit the logistic regression model model <- glm(formula, data = split_train, family = binomial) # Return the model summary (this becomes the list element) summary(model) }) # Name list elements so you can easily reference specific results names(univariate_results) <- predictors
Option 2: Separate Reusable Function
If you want to reuse this logic later, define a dedicated function first:
# Define a function to run univariate logistic regression run_univariate_logistic <- function(var, data) { formula <- as.formula(paste("churn_flag ~", var)) model <- glm(formula, data = data, family = binomial) return(summary(model)) } # Apply the function to all predictors univariate_results <- lapply(predictors, run_univariate_logistic, data = split_train) names(univariate_results) <- predictors
Quick Checks to Avoid Further Issues
- Ensure
churn_flagis binary: Thebinomialfamily inglmrequires the response to be either a factor with two levels (e.g., "Yes"/"No") or a numeric vector of 0s and 1s. If yourchurn_flagis coded differently, convert it first:split_train$churn_flag <- factor(split_train$churn_flag) - Handle missing data:
glmdrops rows with missing values by default. If you need to keep those rows (or want explicit handling), addna.action = na.excludeto yourglmcall. - Access results easily: You can now pull specific summaries using variable names, like
univariate_results[["A1"]]to view the output for predictor A1.
内容的提问来源于stack exchange,提问作者James
相关产品推荐
相关产品推荐

