循环从population data frame创建多data frame并批量提取逻辑回归beta₀
Solution for Extracting Beta₀ from Grouped Logistic Regressions
Got it, let's walk through how to solve this in R—whether you prefer the tidyverse workflow (clean, readable) or base R (no extra packages needed), I've got you covered.
Tidyverse Approach (Recommended for Readability)
First, make sure you have the tidyverse package installed (run install.packages("tidyverse") if you haven't already). This method uses dplyr for grouping and purrr for applying functions to each group:
library(tidyverse) # Replace `your_formula_here` with your actual logistic regression formula # Example: if your outcome is `binary_outcome` and predictors are `var1` + `var2`, use binary_outcome ~ var1 + var2 beta0_results <- pop %>% group_by(replicate) %>% nest() %>% # Nest each replicate's data into a list column mutate( # Fit logistic regression for each nested dataset model = map(data, ~ glm(your_formula_here, data = ., family = binomial(link = "logit"))), # Extract the intercept (beta₀) from each model beta0 = map_dbl(model, ~ coef(.)["(Intercept)"]) ) %>% # Keep only the replicate ID and beta₀ values select(replicate, beta0) %>% ungroup() %>% # Sort by replicate to ensure order 1-110 arrange(replicate)
Quick Explanation:
nest(): Packages all rows for eachreplicateinto its own mini-data frame, making it easy to process groups individually.map(): Applies theglm()function to each nested data frame to fit the logistic regression.map_dbl(): Pulls out the intercept (labeled(Intercept)in model coefficients) as a numeric value, so we end up with a clean column of beta₀s.
Base R Approach
If you'd rather avoid extra packages, here's how to do it with base R functions:
# Split the original data frame into a list of sub-data frames by replicate split_pop <- split(pop, pop$replicate) # Define a helper function to fit the model and extract beta₀ get_beta0 <- function(df) { # Again, replace `your_formula_here` with your actual formula model <- glm(your_formula_here, data = df, family = binomial) # Return the intercept value coef(model)["(Intercept)"] } # Apply the helper function to all sub-data frames beta0_values <- sapply(split_pop, get_beta0) # Convert results to a sorted data frame beta0_results <- data.frame( replicate = as.integer(names(beta0_values)), beta0 = as.numeric(beta0_values) ) %>% arrange(replicate)
Important Notes
- Replace the Formula: Don't forget to swap
your_formula_herewith your actual logistic regression formula. Your outcome variable must be binary (0/1 or TRUE/FALSE) for logistic regression to work. - Handle Small Groups: If some
replicategroups have too few observations (fewer than the number of predictors + 1),glm()will throw an error. To avoid breaking the whole process, usepossibly()in the tidyverse approach to return NA for problematic groups:model = map(data, ~ possibly(glm, otherwise = NA)(your_formula_here, data = ., family = binomial))
内容的提问来源于stack exchange,提问作者wild west
相关产品推荐
相关产品推荐

