R语言:构建模型解释面板数据中行业组间FDI差异的方法
Hey there! Let's walk through how to build your panel data model in R step by step—since you're aiming to explain FDI differences across industry groups using both industry-specific and non-industry-specific variables, we'll cover everything from data prep to model selection and diagnostics.
First, install and load the packages you'll need. plm is the go-to for panel data analysis, dplyr helps with data cleaning, and lmtest/sandwich handle robust standard errors and model tests:
install.packages(c("plm", "dplyr", "lmtest", "sandwich")) library(plm) library(dplyr) library(lmtest) library(sandwich)
Make sure your data has two key identifiers: an industry group ID (like industry_id) and a time column (like year). This tells R it's dealing with panel data.
Check for missing values first—you don't want those messing up your model:
# Load your data (replace with your actual file path) panel_raw <- read.csv("your_fdi_data.csv") # Find rows with missing values missing_rows <- panel_raw %>% filter_all(any_vars(is.na(.))) print(missing_rows)Handle missing data by imputing values or dropping rows (just be cautious with dropping—only do it if the missingness is random).
Convert your data to a panel-specific object so R recognizes its structure:
panel_data <- pdata.frame(panel_raw, index = c("industry_id", "year"))
Now let's test different panel models, starting with basics and moving to more robust options. Your core formula will be:FDI ~ RW + TFP + IY + CY + GDP + LP + IR + ER + PS
3.1 Pooled OLS (Baseline)
This treats all observations as independent, which is simple but ignores industry-specific differences. We'll add clustered standard errors to account for industry-level correlation:
pooled_model <- plm(FDI ~ RW + TFP + IY + CY + GDP + LP + IR + ER + PS, data = panel_data, model = "pooling") # Get clustered robust standard errors (cluster by industry) pooled_clustered <- coeftest(pooled_model, vcov = vcovHC(pooled_model, cluster = "group")) summary(pooled_clustered)
3.2 Fixed Effects (FE) Model
This controls for unobserved, time-invariant industry characteristics (like a sector's inherent attractiveness to FDI). Important note: If non-industry variables (IR, ER, PS) are identical across industries in the same year (e.g., national interest rates), the FE model will drop them—because there's no within-industry variation to estimate.
To work around this, add interaction terms between macro variables and industry-specific traits. This lets you estimate how different industries respond to macro changes:
# Basic FE model (may drop non-industry variables if they have no within-industry variation) fe_model <- plm(FDI ~ RW + TFP + IY + CY + GDP + LP + IR + ER + PS, data = panel_data, model = "within") summary(fe_model) # FE model with interaction terms (captures industry-specific responses to macro variables) fe_interact_model <- plm(FDI ~ RW + TFP + IY + CY + GDP + LP + IR*RW + ER*TFP + PS*IY, data = panel_data, model = "within") summary(fe_interact_model)
3.3 Random Effects (RE) Model
This assumes unobserved industry effects are random rather than fixed. It lets you include time-invariant variables, but you need to test if this assumption holds:
re_model <- plm(FDI ~ RW + TFP + IY + CY + GDP + LP + IR + ER + PS, data = panel_data, model = "random") summary(re_model)
3.4 Hausman Test (Choose FE vs RE)
This test tells you whether fixed effects or random effects is the better fit. A significant p-value means you should go with FE:
hausman_test <- phtest(fe_model, re_model) print(hausman_test)
Don't skip this—you need to ensure your results are reliable:
- Heteroskedasticity check: Use the Breusch-Pagan test. If significant, use robust standard errors:
bptest(fe_model) # Test for heteroskedasticity fe_robust <- coeftest(fe_model, vcov = vcovHC(fe_model, type = "HC1")) # Robust SE - Serial correlation check: Use the Breusch-Godfrey test. If present, use HAC robust standard errors:
pbgtest(fe_model) # Test for serial correlation fe_hac <- coeftest(fe_model, vcov = vcovHAC(fe_model)) # HAC SE
- If non-industry variables are purely macro-level (no variation across industries), their direct effect can't be estimated in FE models—but interaction terms are a great alternative to show how industries react differently to macro changes.
- If your panel is unbalanced (some industries missing years),
plmhandles this automatically, but double-check that missingness isn't biased (e.g., no systematic drop-out of low-FDI industries). - Consider adding a time trend (e.g.,
+ year) to control for linear changes over time that affect all industries.
内容的提问来源于stack exchange,提问作者Restemayer

