Stata转R技术咨询:tab/gen等效语法及回归代码实现
Hey there! Let's walk through how to replicate your Stata workflow in R, step by step. I'll cover the dummy variable creation (your tab ... gen() lines) first, then the regression with lagged variables and clustered standard errors.
tab year, gen(yr) and tab code, gen(cd) Stata's tab ... gen() creates dummy variables for every level of your categorical variable. In R, you have a few straightforward options to do this:
Option 1: Use fastDummies (most similar to Stata's syntax)
This package makes dummy creation super intuitive, just like Stata. First, install and load it along with other essential packages:
# Install packages if you haven't install.packages(c("tidyverse", "fastDummies", "plm", "lmtest", "sandwich")) # Load required packages library(tidyverse) library(fastDummies) library(plm) library(lmtest) library(sandwich)
Then generate your year and code dummies:
# Assuming your dataset is stored in a data frame called `df` df <- df %>% # Generate year dummies with prefix "yr" dummy_cols(select_columns = "year_numeric", prefix = "yr") %>% # Generate code dummies with prefix "cd" dummy_cols(select_columns = "code_numeric", prefix = "cd")
Option 2: Base R with model.matrix() (no extra packages)
If you prefer sticking to base R, use model.matrix() to generate dummies. We use -1 to avoid dropping a reference group (matching Stata's tab ... gen() which creates all levels):
# Create year dummies year_dummies <- model.matrix(~ factor(year_numeric) - 1, data = df) colnames(year_dummies) <- paste0("yr", unique(df$year_numeric)) # Create code dummies code_dummies <- model.matrix(~ factor(code_numeric) - 1, data = df) colnames(code_dummies) <- paste0("cd", unique(df$code_numeric)) # Bind dummies back to your original data frame df <- cbind(df, year_dummies, code_dummies)
tsset) Your Stata tsset code_numeric year_numeric declares the data as panel data. In R, the plm package handles this cleanly:
# Convert to a panel data frame p_df <- pdata.frame(df, index = c("code_numeric", "year_numeric"))
Your Stata regression includes lagged variables, all year/code dummies, a sample filter, and clustered standard errors by code. Here's how to replicate this:
First, filter your sample to only include rows where sample == 1:
p_df_filtered <- p_df %>% filter(sample == 1)
Then run the regression. We use plm() for panel-aware lagged variables, and vcovCL() + coeftest() to get clustered standard errors:
# Fit the model (pooled OLS, matching Stata's `reg` command) model <- plm( fhpolrigaug ~ lag(fhpolrigaug) + lag(lrgdpch) + starts_with("yr") + starts_with("cd"), data = p_df_filtered, model = "pooling" ) # Calculate standard errors clustered by code_numeric clustered_se <- vcovCL(model, cluster = "group") # Display results with clustered SE coeftest(model, vcov = clustered_se)
Alternative: Without plm (using dplyr for lags)
If you'd rather avoid the plm package, you can manually create lagged variables with dplyr:
# Create lagged variables grouped by code_numeric df_filtered <- df %>% filter(sample == 1) %>% group_by(code_numeric) %>% mutate( lag_fhpolrigaug = lag(fhpolrigaug), lag_lrgdpch = lag(lrgdpch) ) %>% ungroup() # Fit OLS regression lm_model <- lm( fhpolrigaug ~ lag_fhpolrigaug + lag_lrgdpch + starts_with("yr") + starts_with("cd"), data = df_filtered ) # Get clustered SE by code_numeric clustered_se_lm <- vcovCL(lm_model, cluster = ~ code_numeric) # Show results coeftest(lm_model, vcov = clustered_se_lm)
Key Notes
starts_with("yr")andstarts_with("cd")replicate Stata'syr*andcd*syntax, selecting all dummy variables starting with those prefixes.- The
lag()function (either inplmordplyr) matches Stata'sL.operator, ensuring you get lagged values within each code group. - Using
vcovCL()gives you the same clustered standard errors as Stata'scluster(code)option.
内容的提问来源于stack exchange,提问作者w12345678

