如何在回归模型中纳入数据框全部变量及其组内一阶滞后项
Hey there! I totally get how tedious it would be to manually create lagged variables for 150+ columns, especially when you need to restrict lags within each group. Let’s walk through two efficient approaches to solve this:
Approach 1: Use dplyr::across to Batch Create Grouped Lags
This method generates all your lagged variables upfront in the data frame, making it easy to inspect and use with standard regression functions like lm().
First, load the dplyr package and define which variables need lagging (exclude group and your response y):
library(dplyr) # Identify variables to lag (all except group and y) vars_to_lag <- setdiff(names(df), c("group", "y")) # Create grouped first-order lags, with clear naming (e.g., x1_lag1) df_with_lags <- df %>% group_by(group) %>% mutate(across(all_of(vars_to_lag), ~lag(.), .names = "{.col}_lag1")) %>% ungroup()
The resulting data frame will have your original variables plus their grouped lags (NAs for the first row of each group, which is correct):
group y x1 x2 x1_lag1 x2_lag1 <chr> <int> <int> <int> <int> <int> 1 A 11 21 1 NA NA 2 A 12 22 2 21 1 3 A 13 23 3 22 2 4 A 14 24 4 23 3 5 B 15 25 5 NA NA 6 B 16 26 6 25 5 7 B 17 27 7 26 6 8 B 18 28 8 27 7
Next, build your regression formula dynamically and fit the model:
# Build the formula string: y ~ x1 + x2 + x1_lag1 + x2_lag1 formula_str <- paste("y ~", paste(c(vars_to_lag, paste0(vars_to_lag, "_lag1")), collapse = " + ")) # Fit the linear model model <- lm(as.formula(formula_str), data = df_with_lags) # Check results summary(model)
Approach 2: Use plm for Panel Data Regression (No Preprocessing Needed)
If you don’t want to clutter your data frame with extra columns, the plm package is perfect for this. It’s designed for panel data and automatically handles grouped lags in the formula.
First, install and load plm, then convert your data to a panel data object (specify group as the individual identifier; if you have a time column, replace "row.names" with that column name):
install.packages("plm") library(plm) # Convert to panel data frame pdata <- pdata.frame(df, index = c("group", "row.names"))
Now you can directly use lag() in your formula—plm will only lag values within each group, no cross-group leakage:
# Fit the model with original variables and their grouped lags model_plm <- plm(y ~ x1 + x2 + lag(x1) + lag(x2), data = pdata, model = "pooling") # View summary summary(model_plm)
For 150+ variables, you can build the formula dynamically to avoid typing every variable:
# Generate formula string automatically formula_plm_str <- paste("y ~", paste(paste0(c(vars_to_lag, paste0("lag(", vars_to_lag, ")"))), collapse = " + ")) # Fit the model model_plm <- plm(as.formula(formula_plm_str), data = pdata, model = "pooling")
Quick Comparison of Approaches
- dplyr + lm: Great if you need to visualize or manipulate the lagged variables before regression. Adds columns to your data frame, which is fine unless you’re working with extremely large datasets.
- plm: More streamlined for panel data scenarios—no need to modify your original data. The
lag()function works natively within groups, making it ideal for large numbers of variables.
内容的提问来源于stack exchange,提问作者martinim

