You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Purrr批量按年份拟合多因变量线性模型并提取参数

批量按年份拟合多因变量线性模型(tidymodels + purrr实现)

问题背景

我有一个可复现的面板数据集,包含location、year以及变量w、x、y、z。需要实现:

  • 以w和x作为因变量,y和z作为解释变量
  • 按年份分别拟合线性模型
  • 汇总所有年份、所有因变量的参数估计结果,形成统一的数据框

现有代码只能手动单个年份、单个因变量拟合,希望用purrr实现自动化批量处理。

可复现数据集

set.seed(123)
df <- tibble(
  location= c(rep("A",5),rep("B",5), rep("C",5), 
              rep("D",5), rep("E",5), rep("F",5),
              rep("G",5), rep("H",5), rep("I",5),
              rep("J",5)),
  year = rep(c(2017,2018,2019,2020,2021),10),
  w = rnorm(50),  
  x = rnorm(50),
  y = rnorm(50),
  z = rnorm(50)
)

现有手动拟合代码

library(tidyverse)
library(tidymodels)

dependent_vars<- c("w","x")
model_1 <- reformulate(termlabels = c("y","z"), response = dependent_vars[1] )

lm_mod <- linear_reg()

lm_fit_2017_w <- 
  lm_mod %>% 
  fit(model_1, data = df %>% filter (year==2017)) %>%
  tidy() %>% 
  mutate(year=2017, response= dependent_vars[1])

# 需重复上述步骤处理其他年份和因变量,最后合并结果

批量拟合解决方案

以下提供两种高效实现方式,核心利用purrr的迭代功能结合tidymodels完成自动化处理:

方法1:交叉组合+嵌套数据框(直观易读)

这种方式逻辑清晰,便于后续扩展或调整:

library(tidyverse)
library(tidymodels)

# 1. 生成所有年份-因变量的组合
param_grid <- crossing(
  year = unique(df$year),
  response_var = c("w", "x")
)

# 2. 按年份嵌套数据,关联参数组合
model_data <- df %>%
  group_by(year) %>%
  nest() %>%
  right_join(param_grid, by = "year")

# 3. 定义拟合函数:输入数据和因变量,返回整理后的参数
fit_lm <- function(data, response) {
  model_formula <- reformulate(termlabels = c("y", "z"), response = response)
  linear_reg() %>%
    fit(model_formula, data = data) %>%
    tidy() %>%
    mutate(response = response)
}

# 4. 批量拟合并展开结果
final_results <- model_data %>%
  mutate(model_results = map2(data, response_var, fit_lm)) %>%
  unnest(model_results) %>%
  select(-data) %>%
  arrange(response, year, term)

方法2:紧凑管道式实现

如果偏好更简洁的代码,可以用expand_grid结合map2直接完成:

library(tidyverse)
library(tidymodels)

dependent_vars <- c("w", "x")
years <- unique(df$year)

# 批量生成组合、拟合模型并整理结果
final_results <- expand_grid(year = years, response = dependent_vars) %>%
  mutate(
    model = map2(year, response, function(y, resp) {
      df %>%
        filter(year == y) %>%
        fit(linear_reg(), reformulate(c("y", "z"), resp), data = .) %>%
        tidy()
    })
  ) %>%
  unnest(model) %>%
  arrange(response, year, term)

结果说明

最终的final_results数据框包含以下核心字段:

  • year:模型拟合的年份
  • response:对应的因变量(w或x)
  • term:模型参数(截距、y、z)
  • estimate:参数估计值
  • std.error:标准误
  • statistic:t统计量
  • p.value:p值

可直接用于后续参数分布分析,比如按因变量和参数分组,查看跨年份的参数变化趋势。

内容的提问来源于stack exchange,提问作者አብርሽ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 20:50:24