You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言合并两个dataframe:ID取并集,二元指标按任一表为1则取1处理

合并逻辑说明

你要实现的是两个患者表的全并集合并,二元指标取「只要任意一个表记录为阳性就标记为阳性」的逻辑,本质可以用any()函数直接实现,完全匹配你的规则:只要存在1就返回1,其余所有情况(全0、全缺失、0+缺失)都返回0。

下面是三种主流R框架的实现方案:

1. tidyverse 实现(可读性最高)

适合中小规模数据,代码逻辑易懂:

library(dplyr)

result <- bind_rows(A, B) %>%
  group_by(id) %>%
  summarise(
    cancer = as.integer(any(cancer == 1, na.rm = TRUE)),
    hiv = as.integer(any(hiv == 1, na.rm = TRUE))
  ) %>%
  ungroup()

2. data.table 实现(性能最优)

适合十万级以上的大样本医疗数据,运行速度是tidyverse的数倍到数十倍:

library(data.table)
# 转换为data.table格式
setDT(A)
setDT(B)

result <- rbind(A, B)[, .(
  cancer = as.integer(any(cancer == 1, na.rm = TRUE)),
  hiv = as.integer(any(hiv == 1, na.rm = TRUE))
), by = id]

3. Base R 实现(无外部依赖)

不需要加载任何第三方包,适合极简环境使用:

# 取两个表id的并集
all_id <- union(A$id, B$id)
result <- data.frame(id = all_id, cancer = 0, hiv = 0)

for (i in seq_len(nrow(result))) {
  current_id <- result$id[i]
  # 提取两个表中当前id的对应记录
  a_vals <- A[A$id == current_id, c("cancer", "hiv")]
  b_vals <- B[B$id == current_id, c("cancer", "hiv")]
  all_vals <- rbind(a_vals, b_vals)
  # 计算两个指标的取值
  result$cancer[i] <- as.integer(any(all_vals$cancer == 1, na.rm = TRUE))
  result$hiv[i] <- as.integer(any(all_vals$hiv == 1, na.rm = TRUE))
}

内容的提问来源于stack exchange,提问作者elbord77

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 03:36:04