如何在R中获取bootstrap模型或稳健标准误模型的标准化回归系数
Great question! I’ve run into this exact issue before—lm.beta is super handy for plain lm objects, but it falls flat when you’re working with bootstrapped models or robust standard errors. The good news is there’s a straightforward fix rooted in what standardized coefficients actually are: they’re just the coefficients you get when you fit your model on z-scored versions of all your variables (dependent and independent alike). Let’s walk through how to apply this to both your use cases:
1. Standardized Coefficients with Robust Standard Errors (sandwich Package)
Instead of trying to retroactively standardize coefficients from your original model, pre-process your data to use z-scored variables first. This way, the model outputs are already standardized, and you can apply robust SE calculations directly:
- First, create a standardized subset of your data (we’ll use
scale()to z-score variables, then convert back to a data frame):
# Standardize only the variables used in your model iris_std <- as.data.frame(scale(iris[, c("Sepal.Length", "Sepal.Width", "Petal.Width")]))
- Fit your linear model using the standardized data:
fit_std <- lm(Sepal.Length ~ Sepal.Width + Petal.Width, data = iris_std)
- Now use
sandwichandlmtestto calculate robust standard errors. The coefficients here are your standardized coefficients, and the SEs are the robust estimates for them:
library(sandwich) library(lmtest) coeftest(fit_std, vcov = vcovHC(fit_std))
2. Standardized Coefficients with Bootstrap (simpleboot Package)
Same core idea: bootstrap the model fitted on standardized variables, and the resulting bootstrapped coefficients will be your standardized coefficients:
- Start with the same standardized dataset as above:
iris_std <- as.data.frame(scale(iris[, c("Sepal.Length", "Sepal.Width", "Petal.Width")]))
- Fit the model on standardized data:
fit_std <- lm(Sepal.Length ~ Sepal.Width + Petal.Width, data = iris_std)
- Run the bootstrap on this standardized model. The output will include the bootstrapped distribution of your standardized coefficients:
library(simpleboot) # Run 1000 bootstrap iterations (adjust R as needed) fit_boot_std <- lm.boot(fit_std, R = 1000) # Inspect the first few bootstrapped standardized coefficients head(fit_boot_std$t) # Get summary statistics for the bootstrapped coefficients summary(fit_boot_std)
Bonus: Manual Calculation for Existing lm Models
If you already have a fitted lm model and don’t want to refit it, you can compute standardized coefficients manually using this formula:standardized_coefficient = original_coefficient * (standard_deviation_of_independent_variable / standard_deviation_of_dependent_variable)
Example code:
# Original model fit.model1.lm <- lm(Sepal.Length ~ Sepal.Width + Petal.Width, iris) # Extract non-intercept coefficients orig_coefs <- coef(fit.model1.lm)[-1] # Calculate standard deviations sd_dep <- sd(iris$Sepal.Length) sd_ind <- sapply(iris[, c("Sepal.Width", "Petal.Width")], sd) # Compute standardized coefficients std_coefs <- orig_coefs * (sd_ind / sd_dep) std_coefs
Note: This manual method works for coefficients alone, but if you need robust SEs or bootstrapped results for these standardized coefficients, pre-standardizing your data is simpler and less error-prone (you’d have to adjust variance matrices or transform each bootstrapped coefficient otherwise).
内容的提问来源于stack exchange,提问作者J. Doe

