如何在R中逐行检查某列值是否存在于其他指定多列中
问题描述
需要逐行检查指定列的数值是否存在于同一行的其他50-60列中,由于列数过多,无法手动指定列名或使用case_when函数处理。尝试用map()函数实现但未达预期,相关信息如下:
示例数据
df1 <- data.frame(A = c(4, 6,3), B = c(4, 1, 1), C = c(1, 1, 3))
尝试的错误代码
nums <- c(4:59) cols <- c(3) wL$Check_Median <- wL[, cols] %>% map(~.x %in% nums) %>% reduce(`|`)
期望逻辑(列名示意)
nums <- c(B:C) cols <- c(A) wL$D <- wL[, cols] %>% map(~.x %in% nums) %>% reduce(`|`)
期望输出
df2 <- data.frame(A = c(4, 6,3), B = c(4, 1, 1), C = c(1, 1, 3), D = c(TRUE, FALSE, TRUE))
解决方案
以下是几种适配多列场景的高效实现方法:
方法1:dplyr行级处理
利用rowwise()和c_across()实现行级判断,适合tidyverse用户:
library(dplyr) # 定义目标列(比如示例中的A列,索引为1) target_col_idx <- 1 # 自动获取除目标列外的所有列索引 other_cols_idx <- setdiff(1:ncol(df1), target_col_idx) df1 <- df1 %>% rowwise() %>% mutate(D = any(c_across(all_of(other_cols_idx)) == !!sym(colnames(df1)[target_col_idx]))) %>% ungroup()
方法2:base R原生实现
无需额外依赖包,执行效率较高:
target_col_idx <- 1 other_cols_idx <- setdiff(1:ncol(df1), target_col_idx) df1$D <- apply(df1, 1, function(row) { row[target_col_idx] %in% row[other_cols_idx] })
方法3:purrr函数式处理
保持tidyverse风格的函数式编程:
library(purrr) library(dplyr) target_col_idx <- 1 other_cols_idx <- setdiff(1:ncol(df1), target_col_idx) df1 <- df1 %>% mutate(D = pmap_lgl(., function(...) { current_row <- c(...) current_row[target_col_idx] %in% current_row[other_cols_idx] }))
核心思路
- 通过
setdiff动态生成其他列的索引,避免手动输入大量列名 - 逐行检查目标列值是否存在于该行的其他列集合中
- 最终生成的
D列即为每行的判断结果
内容的提问来源于stack exchange,提问作者gered
相关产品推荐
相关产品推荐

