You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用MatchIt包结合多数据集,生成匹配2019年特征的数据集

解决方案:匹配2020-2022样本至2019年特征分布

核心思路

以2019年样本为特征参考基准,对2020、2021、2022年的样本进行匹配筛选,保留与2019年特征分布一致的样本,最终合并成目标数据集。以下用MatchIt包实现倾向性得分匹配(PSM),也可根据需求切换精确匹配、卡尺匹配等方式。

步骤1:数据预处理与合并

先给每个年度数据集标记年份,再合并成统一数据集,同时新增标识字段区分基准组(2019年)和待匹配组(2020-2022年)。

假设你的年度数据集分别为df_2019、df_2020、df_2021、df_2022,代码如下:

library(dplyr)
library(MatchIt)

# 标记年份与基准组标识
df_2019 <- df_2019 %>% mutate(year = 2019, is_reference = 1)
df_2020 <- df_2020 %>% mutate(year = 2020, is_reference = 0)
df_2021 <- df_2021 %>% mutate(year = 2021, is_reference = 0)
df_2022 <- df_2022 %>% mutate(year = 2022, is_reference = 0)

# 合并所有数据集(确保各数据集列名、数据类型一致,否则先统一)
combined_df <- bind_rows(df_2019, df_2020, df_2021, df_2022)

步骤2:执行匹配

以2019年样本为基准,匹配后三年样本,匹配变量替换为你关注的特征(如性别、母亲年龄等)。这里采用最近邻匹配,可按需调整参数:

# 执行倾向性得分匹配
match_result <- matchit(is_reference ~ gender + mother_age + child_age + other_features,
                        data = combined_df,
                        method = "nearest", # 最近邻匹配,可选"exact"做精确匹配
                        ratio = 1, # 每个基准样本匹配1个待匹配样本,可调整为2/3等
                        replace = FALSE) # 禁止重复匹配样本

如果需要分年度独立匹配(2020匹配2019,2021匹配2019,2022匹配2019),用循环处理更灵活:

# 初始化空列表存储匹配结果
matched_datasets <- list()

# 分年度循环匹配
for (year in c(2020, 2021, 2022)) {
  # 提取当前年度与2019年数据
  temp_data <- bind_rows(df_2019, get(paste0("df_", year))) %>%
    mutate(is_reference = ifelse(year == 2019, 1, 0))
  
  # 执行匹配
  temp_match <- matchit(is_reference ~ gender + mother_age + other_features,
                        data = temp_data,
                        method = "nearest",
                        ratio = 1)
  
  # 提取当前年度的匹配样本,移除匹配辅助列
  matched_year_data <- match.data(temp_match) %>%
    filter(is_reference == 0) %>%
    select(-is_reference, -distance, -weights)
  
  matched_datasets[[as.character(year)]] <- matched_year_data
}

# 合并匹配后的后三年样本,可选择是否加入2019原样本
final_matched_df <- bind_rows(matched_datasets)
# 若需包含2019样本:final_matched_df <- bind_rows(df_2019, matched_datasets)

步骤3:验证匹配效果

匹配完成后,检查特征分布是否与2019年一致:

# 查看匹配前后的特征平衡统计
summary(match_result)

# 可视化特征分布对比(以母亲年龄为例)
library(ggplot2)
# 匹配前分布
combined_df %>%
  mutate(group = ifelse(is_reference == 1, "2019基准", "原待匹配样本")) %>%
  ggplot(aes(x = mother_age, fill = group)) +
  geom_histogram(alpha = 0.5, position = "identity") +
  facet_wrap(~year)

# 匹配后分布
match.data(match_result) %>%
  mutate(group = ifelse(is_reference == 1, "2019基准", "匹配后样本")) %>%
  ggplot(aes(x = mother_age, fill = group)) +
  geom_histogram(alpha = 0.5, position = "identity") +
  facet_wrap(~year)

常见问题修复

  • 若之前bind_rows失败,先检查各年度数据集的列名、数据类型是否统一,用select保留共同列,mutate转换数据类型。
  • 匹配后样本量过少时,可调大ratio参数(如设为2),或添加caliper = 0.2放宽匹配卡尺。
  • 分类特征优先用method = "exact"精确匹配,连续特征用倾向性得分匹配更高效。

内容的提问来源于stack exchange,提问作者Luis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 23:43:08