You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中检查多个数据框的公共ID并生成统计矩阵

多DataFrame公共ID检查与矩阵生成

需求说明

拥有多个共享ID列的DataFrame,需检查任意两两之间是否存在公共ID,生成矩阵:无公共ID的位置值为0;同时合并所有DataFrame后发现存在重复ID,需定位重复来源。


步骤实现

1. 准备示例数据

dfs <- list(
  DB07 = data.frame(ID = c("A548480901", "A548480902", "A548480903",
                           "A548480904", "A548560901", "A548560902")), 
  DB08 = data.frame(ID = c("A448440901", "A448440902", "A448680102",
                           "A448680501", "A448680502", "A448680503"))
)

2. 提取各DataFrame的唯一ID集合

先提取每个DataFrame的唯一ID,避免单DataFrame内的重复ID干扰跨表判断:

id_sets <- lapply(dfs, function(x) unique(x$ID))

3. 生成公共ID数量矩阵

矩阵中每个位置的值为对应两个DataFrame的公共ID数量,无公共ID则为0:

# 生成对比矩阵
compare_matrix <- outer(names(id_sets), names(id_sets), 
                        function(a, b) sapply(seq_along(a), function(i) {
                          length(intersect(id_sets[[a[i]]], id_sets[[b[i]]]))
                        }))

# 设置行列名
rownames(compare_matrix) <- names(id_sets)
colnames(compare_matrix) <- names(id_sets)

运行后得到的矩阵示例:

DB07 DB08
DB07    6    0
DB08    0    6

4. 可选:生成存在性标记矩阵

若只需标记是否存在公共ID(1=存在,0=不存在),可修改代码:

exist_matrix <- outer(names(id_sets), names(id_sets), 
                      function(a, b) sapply(seq_along(a), function(i) {
                        as.integer(length(intersect(id_sets[[a[i]]], id_sets[[b[i]]])) > 0)
                      }))

rownames(exist_matrix) <- names(id_sets)
colnames(exist_matrix) <- names(id_sets)

对应的结果矩阵:

DB07 DB08
DB07    1    0
DB08    0    1

内容的提问来源于stack exchange,提问作者D. Shin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 15:52:37