You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

调整R语言plm包中plm()函数的多重共线性敏感度

Aligning plm (R) and xtreg (Stata) Results: Fixing Collinearity Differences

Great question—this is a common frustration when moving between R and Stata for panel data analysis. Let's break down your two questions and walk through practical, actionable solutions:

(a) Modifying the Collinearity Threshold in plm

Unfortunately, plm doesn’t have a built-in parameter to directly adjust its collinearity detection threshold (unlike Stata, where you can tweak it with set collinear). But you can manually preprocess your data to match Stata’s threshold rules:

  • Understand Stata’s default threshold: Stata’s xtreg automatically drops variables with a tolerance (1/VIF) less than 1e-7. To replicate this in R:
    1. Use the car package’s vif() function to calculate variance inflation factors for your model variables.
    2. Filter variables to keep only those with a tolerance ≥ 1e-7.
  • Example code:
    library(car)
    library(plm)
    
    # Load sample panel data (replace with your data)
    data("Grunfeld", package = "plm")
    
    # Define your full model formula
    full_formula <- inv ~ value + capital + extra_var1 + extra_var2
    
    # Fit a linear model to compute VIF (works with lm objects)
    lm_temp <- lm(full_formula, data = Grunfeld)
    vif_vals <- vif(lm_temp)
    tolerance_vals <- 1 / vif_vals
    
    # Keep variables matching Stata's default tolerance threshold
    keep_vars <- names(tolerance_vals[tolerance_vals >= 1e-7])
    adjusted_formula <- as.formula(paste("inv ~", paste(keep_vars, collapse = " + ")))
    
    # Run plm with the filtered variables
    plm_aligned <- plm(adjusted_formula, data = Grunfeld, model = "within")
    summary(plm_aligned)
    
  • If you’ve adjusted Stata’s threshold (e.g., set collinear 1e-6), just swap 1e-7 for your custom value in the code above.

(b) Aligning Variable Selection Mechanism with Stata

The bigger difference between plm and xtreg is how they choose which collinear variables to drop:

  • Stata uses the order of variables in your formula: it keeps earlier variables and drops later ones that are linearly dependent.
  • plm relies on a rank-based check of the model matrix, which doesn’t prioritize variable order the same way.

To fix this, you have two options:

Option 1: Manual Alignment (Easiest)

  1. Run your model in Stata first, and note which variables are labeled omitted due to collinearity in the output.
  2. Exact those same variables from your R formula before running plm. For example, if Stata drops extra_var2, your R formula becomes inv ~ value + capital + extra_var1.

Option 2: Automate Stata’s Order-Based Dropping (Advanced)

If you want to replicate Stata’s logic without switching to Stata first, you can build the model matrix incrementally, adding variables one-by-one in formula order and skipping any that introduce collinearity:

library(plm)

data("Grunfeld", package = "plm")
full_formula <- inv ~ value + capital + extra_var1 + extra_var2

# Extract response and predictor variable names
response <- as.character(full_formula[[2]])
predictors <- attr(terms(full_formula), "term.labels")

# Initialize valid variables and matrix
valid_predictors <- c()
current_matrix <- matrix(nrow = nrow(Grunfeld), ncol = 0)

for (var in predictors) {
  # Add current variable to the matrix
  new_col <- Grunfeld[[var]]
  temp_matrix <- cbind(current_matrix, new_col)
  # Check if adding the variable increases the matrix rank (no collinearity)
  if (qr(temp_matrix)$rank == ncol(temp_matrix)) {
    valid_predictors <- c(valid_predictors, var)
    current_matrix <- temp_matrix
  } else {
    cat(paste("Omitting", var, "due to collinearity (matches Stata's logic)\n"))
  }
}

# Build aligned formula and run plm
aligned_formula <- as.formula(paste(response, "~", paste(valid_predictors, collapse = " + ")))
plm_aligned <- plm(aligned_formula, data = Grunfeld, model = "within")
summary(plm_aligned)

Critical Checks for Full Alignment

  • Match model types: Ensure plm’s model parameter matches Stata’s specification:
    • plm(..., model = "within") ↔ xtreg ..., fe
    • plm(..., model = "random") ↔ xtreg ..., re
  • Fixed effects consistency: If using time/individual fixed effects, make sure they’re included identically (e.g., plm(..., effect = "twoways") matches Stata’s xtreg ..., fe i.year i.firm).
  • Missing values: Both tools use listwise deletion, but double-check that your data has no hidden differences in missing value handling.

内容的提问来源于stack exchange,提问作者Durden

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 04:26:57