调整R语言plm包中plm()函数的多重共线性敏感度
plm (R) and xtreg (Stata) Results: Fixing Collinearity Differences Great question—this is a common frustration when moving between R and Stata for panel data analysis. Let's break down your two questions and walk through practical, actionable solutions:
(a) Modifying the Collinearity Threshold in plm
Unfortunately, plm doesn’t have a built-in parameter to directly adjust its collinearity detection threshold (unlike Stata, where you can tweak it with set collinear). But you can manually preprocess your data to match Stata’s threshold rules:
- Understand Stata’s default threshold: Stata’s
xtregautomatically drops variables with a tolerance (1/VIF) less than1e-7. To replicate this in R:- Use the
carpackage’svif()function to calculate variance inflation factors for your model variables. - Filter variables to keep only those with a tolerance ≥
1e-7.
- Use the
- Example code:
library(car) library(plm) # Load sample panel data (replace with your data) data("Grunfeld", package = "plm") # Define your full model formula full_formula <- inv ~ value + capital + extra_var1 + extra_var2 # Fit a linear model to compute VIF (works with lm objects) lm_temp <- lm(full_formula, data = Grunfeld) vif_vals <- vif(lm_temp) tolerance_vals <- 1 / vif_vals # Keep variables matching Stata's default tolerance threshold keep_vars <- names(tolerance_vals[tolerance_vals >= 1e-7]) adjusted_formula <- as.formula(paste("inv ~", paste(keep_vars, collapse = " + "))) # Run plm with the filtered variables plm_aligned <- plm(adjusted_formula, data = Grunfeld, model = "within") summary(plm_aligned) - If you’ve adjusted Stata’s threshold (e.g.,
set collinear 1e-6), just swap1e-7for your custom value in the code above.
(b) Aligning Variable Selection Mechanism with Stata
The bigger difference between plm and xtreg is how they choose which collinear variables to drop:
- Stata uses the order of variables in your formula: it keeps earlier variables and drops later ones that are linearly dependent.
plmrelies on a rank-based check of the model matrix, which doesn’t prioritize variable order the same way.
To fix this, you have two options:
Option 1: Manual Alignment (Easiest)
- Run your model in Stata first, and note which variables are labeled
omitted due to collinearityin the output. - Exact those same variables from your R formula before running
plm. For example, if Stata dropsextra_var2, your R formula becomesinv ~ value + capital + extra_var1.
Option 2: Automate Stata’s Order-Based Dropping (Advanced)
If you want to replicate Stata’s logic without switching to Stata first, you can build the model matrix incrementally, adding variables one-by-one in formula order and skipping any that introduce collinearity:
library(plm) data("Grunfeld", package = "plm") full_formula <- inv ~ value + capital + extra_var1 + extra_var2 # Extract response and predictor variable names response <- as.character(full_formula[[2]]) predictors <- attr(terms(full_formula), "term.labels") # Initialize valid variables and matrix valid_predictors <- c() current_matrix <- matrix(nrow = nrow(Grunfeld), ncol = 0) for (var in predictors) { # Add current variable to the matrix new_col <- Grunfeld[[var]] temp_matrix <- cbind(current_matrix, new_col) # Check if adding the variable increases the matrix rank (no collinearity) if (qr(temp_matrix)$rank == ncol(temp_matrix)) { valid_predictors <- c(valid_predictors, var) current_matrix <- temp_matrix } else { cat(paste("Omitting", var, "due to collinearity (matches Stata's logic)\n")) } } # Build aligned formula and run plm aligned_formula <- as.formula(paste(response, "~", paste(valid_predictors, collapse = " + "))) plm_aligned <- plm(aligned_formula, data = Grunfeld, model = "within") summary(plm_aligned)
Critical Checks for Full Alignment
- Match model types: Ensure
plm’smodelparameter matches Stata’s specification:plm(..., model = "within")↔xtreg ..., feplm(..., model = "random")↔xtreg ..., re
- Fixed effects consistency: If using time/individual fixed effects, make sure they’re included identically (e.g.,
plm(..., effect = "twoways")matches Stata’sxtreg ..., fe i.year i.firm). - Missing values: Both tools use listwise deletion, but double-check that your data has no hidden differences in missing value handling.
内容的提问来源于stack exchange,提问作者Durden

