如何用Purrr批量按年份拟合多因变量线性模型并提取参数
批量按年份拟合多因变量线性模型(tidymodels + purrr实现)
问题背景
我有一个可复现的面板数据集,包含location、year以及变量w、x、y、z。需要实现:
- 以
w和x作为因变量,y和z作为解释变量 - 按年份分别拟合线性模型
- 汇总所有年份、所有因变量的参数估计结果,形成统一的数据框
现有代码只能手动单个年份、单个因变量拟合,希望用purrr实现自动化批量处理。
可复现数据集
set.seed(123) df <- tibble( location= c(rep("A",5),rep("B",5), rep("C",5), rep("D",5), rep("E",5), rep("F",5), rep("G",5), rep("H",5), rep("I",5), rep("J",5)), year = rep(c(2017,2018,2019,2020,2021),10), w = rnorm(50), x = rnorm(50), y = rnorm(50), z = rnorm(50) )
现有手动拟合代码
library(tidyverse) library(tidymodels) dependent_vars<- c("w","x") model_1 <- reformulate(termlabels = c("y","z"), response = dependent_vars[1] ) lm_mod <- linear_reg() lm_fit_2017_w <- lm_mod %>% fit(model_1, data = df %>% filter (year==2017)) %>% tidy() %>% mutate(year=2017, response= dependent_vars[1]) # 需重复上述步骤处理其他年份和因变量,最后合并结果
批量拟合解决方案
以下提供两种高效实现方式,核心利用purrr的迭代功能结合tidymodels完成自动化处理:
方法1:交叉组合+嵌套数据框(直观易读)
这种方式逻辑清晰,便于后续扩展或调整:
library(tidyverse) library(tidymodels) # 1. 生成所有年份-因变量的组合 param_grid <- crossing( year = unique(df$year), response_var = c("w", "x") ) # 2. 按年份嵌套数据,关联参数组合 model_data <- df %>% group_by(year) %>% nest() %>% right_join(param_grid, by = "year") # 3. 定义拟合函数:输入数据和因变量,返回整理后的参数 fit_lm <- function(data, response) { model_formula <- reformulate(termlabels = c("y", "z"), response = response) linear_reg() %>% fit(model_formula, data = data) %>% tidy() %>% mutate(response = response) } # 4. 批量拟合并展开结果 final_results <- model_data %>% mutate(model_results = map2(data, response_var, fit_lm)) %>% unnest(model_results) %>% select(-data) %>% arrange(response, year, term)
方法2:紧凑管道式实现
如果偏好更简洁的代码,可以用expand_grid结合map2直接完成:
library(tidyverse) library(tidymodels) dependent_vars <- c("w", "x") years <- unique(df$year) # 批量生成组合、拟合模型并整理结果 final_results <- expand_grid(year = years, response = dependent_vars) %>% mutate( model = map2(year, response, function(y, resp) { df %>% filter(year == y) %>% fit(linear_reg(), reformulate(c("y", "z"), resp), data = .) %>% tidy() }) ) %>% unnest(model) %>% arrange(response, year, term)
结果说明
最终的final_results数据框包含以下核心字段:
year:模型拟合的年份response:对应的因变量(w或x)term:模型参数(截距、y、z)estimate:参数估计值std.error:标准误statistic:t统计量p.value:p值
可直接用于后续参数分布分析,比如按因变量和参数分组,查看跨年份的参数变化趋势。
内容的提问来源于stack exchange,提问作者አብርሽ
相关产品推荐
相关产品推荐

