如何使用fixest包遍历多变量组合生成回归对象列表?
解决fixest包遍历回归组合的问题
需求说明
需要遍历officer_race、primary_ind、outcome三个向量的所有组合,使用fixest包的feols()运行回归,最终生成包含所有有效回归对象的列表。用户原代码在lm()中可行,但在feols()中失效。
数据集与变量向量
library(fixest) library(tibble) analysis <- tibble( off_race = c("hispanic", "hispanic", "white","white", "hispanic", "white", "hispanic", "white", "white", "white","hispanic"), any_black_uof = c(FALSE, FALSE, FALSE, FALSE, FALSE, FALSE, FALSE, FALSE, FALSE, FALSE, FALSE), any_black_arrest = c(TRUE, FALSE, FALSE, FALSE, FALSE, FALSE, FALSE, FALSE, FALSE, FALSE, FALSE), prop_white_scale = c(0.866619646027524, -1.14647499712298, 1.33793539994219, 0.593565300512359, -0.712819809606193, 0.3473585867755, -1.37025501425243, 1.16596624239715, 0.104521426674564, 0.104521426674564, -1.53728347122581), prop_hisp_scale = c(-0.347382203637802, 1.54966785579018, -0.833021026477168, -0.211470492567308, 1.48353691981021, 0.421968013870802, 2.63739845069911, -0.61002505397242, 0.66674880256898,0.66674880256898, 2.93190487813111) ) officer_race = c("black", "white", "hispanic") primary_ind <- c("prop_white_scale","prop_hisp_scale","prop_black_scale") outcome <- c("any_black_uof","any_white_uof","any_hisp_uof","any_black_arrest","any_white_arrest","any_hisp_arrest","any_black_stop","any_white_stop","any_hisp_stop")
问题分析
原代码失效的核心原因:
- 公式逻辑错误:把因变量和自变量写反了(原代码构造的是
predictor ~ outcome,正确应为outcome ~ predictor),feols()对公式的逻辑校验更严格,直接导致回归逻辑错误。 - 未处理缺失变量/空数据:原数据集没有
prop_black_scale和大部分outcome变量,且officer_race中的black无对应数据,运行时会直接报错中断。
解决方案代码
library(dplyr) # 生成所有参数组合的网格 param_grid <- expand.grid( off_race = officer_race, primary_ind = primary_ind, outcome = outcome, stringsAsFactors = FALSE ) # 遍历每个组合运行回归 result_list <- lapply(1:nrow(param_grid), function(i) { # 获取当前参数 curr_race <- param_grid$off_race[i] curr_pred <- param_grid$primary_ind[i] curr_outcome <- param_grid$outcome[i] # 过滤数据:保留对应种族,且确保变量存在并移除缺失值 filtered_data <- analysis %>% filter(off_race == curr_race) %>% select(all_of(c(curr_outcome, curr_pred))) %>% drop_na() # 空数据则返回NULL并给出警告 if(nrow(filtered_data) == 0) { warning(sprintf("无匹配数据:警官种族=%s,因变量=%s,自变量=%s", curr_race, curr_outcome, curr_pred)) return(NULL) } # 构造正确的回归公式 reg_formula <- reformulate(curr_pred, response = curr_outcome) # 运行fixest回归 feols(reg_formula, data = filtered_data) }) # 给列表元素命名,方便索引 names(result_list) <- apply(param_grid, 1, function(x) { sprintf("off_race=%s_outcome=%s_predictor=%s", x[1], x[3], x[2]) }) # 移除无数据的NULL元素 result_list <- Filter(Negate(is.null), result_list)
代码说明
- 参数网格生成:用
expand.grid()生成所有可能的参数组合,确保不遗漏任何回归场景。 - 数据过滤与校验:先筛选对应种族的数据,再仅保留当前回归需要的变量,同时移除缺失值;空数据时给出警告并跳过,避免程序崩溃。
- 公式构造:用
reformulate()明确指定因变量和自变量,避免公式顺序错误。 - 结果整理:给列表元素命名,方便后续查找特定回归结果;移除无效的NULL元素,只保留有效回归对象。
内容的提问来源于stack exchange,提问作者sd3184
相关产品推荐
相关产品推荐

