R语言实现两个数据集对应行取值一致性对比的更优方法
R按行对比两个数据集对应行取值一致性的实现方案
适用场景
两个数据集行数一致、待对比列名相同,需要逐行判断所有列取值是否完全相等,最终给其中一个数据集新增标识列。
实现方案
方法1:base R 单行实现(效率最高,适合大数据集)
如果两个数据集列顺序完全一致,直接用rowSums对比即可:
df1 <- data.frame(a=c(1,2,3),b=c(4,5,6)) df2 <- data.frame(a=c(1,2,4),b=c(4,5,6)) # 逐行对比所有列是否全等,生成标识列 df2$c <- as.character(rowSums(df1 == df2) == ncol(df1))
运行结果:
> df2 a b c 1 1 4 T 2 2 5 T 3 4 6 F
如果数据集中存在NA值,用兼容NA的写法:
df2$c <- as.character(apply(df1 == df2, 1, \(x) all(x, na.rm = TRUE)))
方法2:dplyr 实现(兼容列顺序不一致场景)
如果两个数据集列顺序可能不同,先对齐公共列再对比:
library(dplyr) # 提取公共待对比列 common_cols <- intersect(colnames(df1), colnames(df2)) final <- df2 %>% mutate( c = rowSums(select(df1, all_of(common_cols)) == select(., all_of(common_cols))) == length(common_cols), c = as.character(c) )
行数不一致场景适配
如果两个数据集行数不同,可先按行索引合并再对比:
library(tibble) library(dplyr) common_cols <- intersect(colnames(df1), colnames(df2)) # 给两个数据集新增行索引列 df1_idx <- rownames_to_column(df1, "row_id") df2_idx <- rownames_to_column(df2, "row_id") final <- full_join(df1_idx, df2_idx, by = "row_id", suffix = c("_df1", "_df2")) %>% mutate( c = rowSums(select(., ends_with("_df1")) == select(., ends_with("_df2")), na.rm = TRUE) == length(common_cols), c = as.character(c) )
内容的提问来源于stack exchange,提问作者younghyun
相关产品推荐
相关产品推荐

