You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何判断列表中各DataFrame的ID是否唯一并生成汇总表

嘿,我来帮你解决这个问题!针对你手里这个包含多个DataFrame的列表,我们可以用R里的工具快速完成ID列唯一性检查,还能生成清晰的汇总表,分两种方案给你参考:

方案一:用tidyverse工具包(推荐,代码更简洁易读)

首先确保你已经安装并加载了tidyverse,它包含了我们需要的dplyr和purrr工具:

library(tidyverse)

先把你的原始列表格式化一下,看起来更清晰:

lst <- list(
  # 第一个DataFrame
  structure(list(ID = c("Tom", "Jerry", "Mary"), Score = c(85, 85, 96)), 
            row.names = c(NA, -3L), class = c("tbl_df", "tbl", "data.frame")),
  # 第二个DataFrame(存在重复ID:Jerry)
  structure(list(ID = c("Tom", "Jerry", "Mary", "Jerry"), Score = c(75, 65, 88, 98)), 
            row.names = c(NA, -4L), class = c("tbl_df", "tbl", "data.frame")),
  # 第三个DataFrame(存在重复ID:Tom)
  structure(list(ID = c("Tom", "Jerry", "Tom"), Score = c(97, 65, 96)), 
            row.names = c(NA, -3L), class = c("tbl_df", "tbl", "data.frame"))
)

然后用imap_dfr遍历列表,同时获取每个DataFrame的索引和内容,一键生成汇总表:

summary_table <- imap_dfr(lst, function(df, idx) {
  tibble(
    dataframe_number = idx,          # 标记是列表中的第几个DataFrame
    id_unique = n_distinct(df$ID) == nrow(df),  # 判断ID是否唯一
    total_records = nrow(df),        # 该DataFrame的总行数
    distinct_id_count = n_distinct(df$ID),  # 不重复的ID数量
    duplicate_id_count = total_records - distinct_id_count  # 重复的ID记录数
  )
})

# 查看最终汇总表
summary_table

运行后会得到这样清晰的结果:

# A tibble: 3 × 5
  dataframe_number id_unique total_records distinct_id_count duplicate_id_count
             <int> <lgl>             <int>             <int>              <int>
1                1 TRUE                  3                 3                  0
2                2 FALSE                 4                 3                  1
3                3 FALSE                 3                 2                  1

方案二:用基础R实现(无需额外安装包)

如果你不想加载tidyverse,用基础R的lapply和do.call也能完成同样的工作:

summary_table_base <- do.call(rbind, lapply(seq_along(lst), function(idx) {
  df <- lst[[idx]]
  total_rows <- nrow(df)
  unique_ids <- length(unique(df$ID))
  
  data.frame(
    dataframe_number = idx,
    id_unique = unique_ids == total_rows,
    total_records = total_rows,
    distinct_id_count = unique_ids,
    duplicate_id_count = total_rows - unique_ids
  )
}))

# 查看结果
summary_table_base

这个结果和方案一完全一致,只是用了基础R的原生函数实现。

关键逻辑说明

判断ID唯一性的核心逻辑很简单:不重复ID的数量是否等于DataFrame的总行数。如果相等,说明每个ID只出现一次;如果不等,两者的差值就是重复的记录数。同时给每个DataFrame加上索引标记,方便你对应回原始列表中的具体数据。

内容的提问来源于stack exchange,提问作者Stataq

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 14:07:29