You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中基于匹配ID列将另一数据集的cohort列合并至目标数据集

将df2的cohort列合并至df1的实现方法

示例数据

先加载你提供的两个数据集:

df1 <- data.frame(
  ID = c(1, 1, 1, 1, 1, 1), 
  rater_type = c("Direct", "Direct", "Other", "Other", "Peer", "Peer"), 
  year = c(1, 2, 1, 2, 1, 2), 
  m1 = c(4, 5, NA, 4.5, 3, 4), 
  m2 = c(4, 5, NA, 3.5, 5, 4)
)

df2 <- data.frame(
  ID = c(1, 1, 1, 1, 1, 1),
  cohort = c(2, 2, 2, 2, 2, 2),
  rater_type = c("Direct", "Direct", "Other", "Other", "Peer", "Peer"),
  competency_new = c("m1", "m1", "m1", "m2", "m2", "m2"),
  year = c(1, 2, 1, 2, 1, 2),
  score = c(4, 5, NA, 3.5, 5, 4)
)

实现方法

方法一:Base R 原生实现

用merge()函数按ID匹配,保留df1的所有行,同时提取df2中唯一的ID-cohort映射(避免重复匹配):

# 提取唯一的ID-cohort对应关系
df2_cohort_map <- unique(df2[, c("ID", "cohort")])
# 合并数据集
df3 <- merge(df1, df2_cohort_map, by = "ID", all.x = TRUE)
# 调整列顺序匹配期望输出
df3 <- df3[, c("ID", "cohort", "rater_type", "year", "m1", "m2")]

方法二:dplyr 包实现(tidyverse 生态)

用left_join()做左连接,结合distinct()确保每个ID只对应一个cohort值,代码更简洁直观:

library(dplyr)

df3 <- df1 %>%
  left_join(distinct(df2, ID, cohort), by = "ID") %>%
  select(ID, cohort, rater_type, year, m1, m2)

验证输出

运行上述代码后,得到的df3与你期望的输出完全一致:

> df3
  ID cohort rater_type year  m1  m2
1  1      2      Direct    1 4.0 4.0
2  1      2      Direct    2 5.0 5.0
3  1      2       Other    1  NA  NA
4  1      2       Other    2 4.5 3.5
5  1      2        Peer    1 3.0 5.0
6  1      2        Peer    2 4.0 4.0

内容的提问来源于stack exchange,提问作者sdS

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 15:45:07