Probit模型中行业与年份固定效应的实现及R模板求助
Probit模型中加入行业与年份固定效应的实现方法及R代码模板
一、等效实现的核心逻辑
Probit作为非线性模型,无法像OLS那样通过组内demean直接吸收固定效应,行业与年份固定效应的常用等效形式是引入行业虚拟变量和年份虚拟变量:
- 行业固定效应:将行业变量转换为因子型,纳入模型后每个行业对应一个虚拟变量(自动保留一个基准组避免多重共线性),用于控制行业层面不随时间变化的异质性(比如行业准入门槛、技术特征等)。
- 年份固定效应:同理,将年份变量转换为因子型纳入模型,控制跨年度的共同冲击(比如宏观经济波动、政策变化等)。
这种方法适用于混合截面数据或短面板数据(时间维度T较小),能有效替代"固定效应"的控制逻辑。
二、R代码模板
方法1:混合Probit + 虚拟变量(通用场景)
适合混合截面或短面板数据,直接通过虚拟变量控制行业和年份固定效应:
# 加载依赖包 library(dplyr) library(stats) # 定义带固定效应的Probit拟合函数 probit_fe_fit <- function(data, dep_var, indep_vars, industry_col, year_col) { # 将行业、年份转换为因子(自动生成虚拟变量) data <- data %>% mutate(across(c({{industry_col}}, {{year_col}}), as.factor)) # 构建模型公式 model_formula <- as.formula( paste0(dep_var, " ~ ", paste(indep_vars, collapse = " + "), " + ", industry_col, " + ", year_col) ) # 拟合Probit模型 fit <- glm(model_formula, data = data, family = binomial(link = "probit")) return(fit) } # ------------------------------ # 示例用法(模拟数据演示) # ------------------------------ set.seed(123) sample_data <- tibble( y = rbinom(1000, 1, 0.3), # 二元因变量 x1 = rnorm(1000), # 核心自变量1 x2 = rnorm(1000), # 核心自变量2 industry = sample(paste0("Ind", 1:10), 1000, replace = TRUE), # 行业分类 year = sample(2010:2020, 1000, replace = TRUE) # 年份 ) # 调用函数拟合模型 model_result <- probit_fe_fit( data = sample_data, dep_var = "y", indep_vars = c("x1", "x2"), industry_col = "industry", year_col = "year" ) # 查看结果 summary(model_result)
方法2:面板Probit + 固定效应(追踪面板场景)
如果是个体/企业追踪的面板数据,需要同时控制个体(或行业)固定效应和年份固定效应,可以使用pglm包实现面板固定效应Probit:
# 加载依赖包 library(pglm) library(dplyr) # 定义面板Probit拟合函数 panel_probit_fe_fit <- function(data, dep_var, indep_vars, id_col, year_col) { # 转换个体、年份为因子 data <- data %>% mutate({{id_col}} := as.factor({{id_col}}), {{year_col}} := as.factor({{year_col}})) # 构建公式:个体固定效应 + 年份虚拟变量 model_formula <- as.formula( paste0(dep_var, " ~ ", paste(indep_vars, collapse = " + "), " + ", year_col, " + factor(", id_col, ")") ) # 拟合固定效应面板Probit模型 fit <- pglm(model_formula, data = data, family = binomial(link = "probit"), model = "fixed", index = c({{id_col}}, {{year_col}})) return(fit) } # ------------------------------ # 示例用法(模拟面板数据演示) # ------------------------------ set.seed(456) panel_data <- tibble( firm_id = rep(paste0("Firm", 1:200), each = 10), # 企业个体ID year = rep(2011:2020, 200), # 年份 y = rbinom(2000, 1, 0.4), # 二元因变量 x1 = rnorm(2000), # 核心自变量1 x2 = rnorm(2000), # 核心自变量2 industry = rep(sample(paste0("Ind", 1:10), 200, replace = TRUE), each = 10) ) # 拟合模型(控制企业固定效应+年份固定效应) panel_model <- panel_probit_fe_fit( data = panel_data, dep_var = "y", indep_vars = c("x1", "x2"), id_col = "firm_id", year_col = "year" ) summary(panel_model)
三、关键注意事项
- 虚拟变量法中,
glm和pglm会自动处理基准组,无需手动删除多余虚拟变量。 - 若行业/年份类别过多,模型参数数量会增加,但样本量足够时不影响拟合效果。
- 长面板(T较大)使用固定效应Probit时,会存在incidental parameter偏误,此时可考虑随机效应模型,或使用
cmplm等包的修正方法。
内容的提问来源于stack exchange,提问作者Robbert
相关产品推荐
相关产品推荐

