You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中逐行检查某列值是否存在于其他指定多列中

问题描述

需要逐行检查指定列的数值是否存在于同一行的其他50-60列中,由于列数过多,无法手动指定列名或使用case_when函数处理。尝试用map()函数实现但未达预期,相关信息如下:

示例数据

df1 <- data.frame(A = c(4, 6,3), 
                  B = c(4, 1, 1), 
                  C = c(1, 1, 3))

尝试的错误代码

nums <- c(4:59)
cols <- c(3)

wL$Check_Median <-
  wL[, cols] %>%
  map(~.x %in% nums) %>%
  reduce(`|`)

期望逻辑(列名示意)

nums <- c(B:C)
cols <- c(A)

wL$D <-
  wL[, cols] %>%
  map(~.x %in% nums) %>%
  reduce(`|`)

期望输出

df2 <- data.frame(A = c(4, 6,3), 
                  B = c(4, 1, 1), 
                  C = c(1, 1, 3),
                  D = c(TRUE, FALSE, TRUE))
解决方案

以下是几种适配多列场景的高效实现方法:

方法1:dplyr行级处理

利用rowwise()和c_across()实现行级判断,适合tidyverse用户:

library(dplyr)

# 定义目标列(比如示例中的A列,索引为1)
target_col_idx <- 1
# 自动获取除目标列外的所有列索引
other_cols_idx <- setdiff(1:ncol(df1), target_col_idx)

df1 <- df1 %>%
  rowwise() %>%
  mutate(D = any(c_across(all_of(other_cols_idx)) == !!sym(colnames(df1)[target_col_idx]))) %>%
  ungroup()

方法2:base R原生实现

无需额外依赖包,执行效率较高:

target_col_idx <- 1
other_cols_idx <- setdiff(1:ncol(df1), target_col_idx)

df1$D <- apply(df1, 1, function(row) {
  row[target_col_idx] %in% row[other_cols_idx]
})

方法3:purrr函数式处理

保持tidyverse风格的函数式编程:

library(purrr)
library(dplyr)

target_col_idx <- 1
other_cols_idx <- setdiff(1:ncol(df1), target_col_idx)

df1 <- df1 %>%
  mutate(D = pmap_lgl(., function(...) {
    current_row <- c(...)
    current_row[target_col_idx] %in% current_row[other_cols_idx]
  }))

核心思路

  • 通过setdiff动态生成其他列的索引,避免手动输入大量列名
  • 逐行检查目标列值是否存在于该行的其他列集合中
  • 最终生成的D列即为每行的判断结果

内容的提问来源于stack exchange,提问作者gered

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 00:35:30