You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何判断DataFrame中前缀为i10_pr的列是否匹配列表值并生成标记列

问题描述

我有一个DataFrame(df1),包含约100个以i10_pr开头的列。需要对每一行检查:这些以i10_pr开头的列中是否有任意值与另一个列表(df2)中的值匹配,并生成新列match,匹配则为1,不匹配则为0。

期望输出如下:

#Desired output
  visitlink visitorder i10_pr1   i10_pr2   i10_pr3   i10_pr4   i10_pr5   match
1   7466851          3 "BW28ZZZ" "BR30Y0Z" "BR39Y0Z" ""        ""        0
2   7023336          1 "0BDC8ZX" "0BDC8ZX" "07D78ZX" ""        ""        0
3   2481935          3 "5A09357" "3C1ZX8Z" "06HN33Z" "B54CZZA" "0W993ZX" 1
4   4605446          1 "5A1955Z" "0BH17EZ" "03HY32Z" "02HV33Z" "B548ZZA" 1
5   7287173          2 ""        ""        ""        ""        ""        0

测试数据:

# df1
structure(list(visitlink = c(7466851, 7023336, 2481935, 4605446, 
7287173), visitorder = c(3L, 1L, 3L, 1L, 2L), i10_pr1 = c("BW28ZZZ", 
"0BDC8ZX", "5A09357", "5A1955Z", ""), i10_pr2 = c("BR30Y0Z", 
"0BDC8ZX", "3C1ZX8Z", "0BH17EZ", ""), i10_pr3 = c("BR39Y0Z", 
"07D78ZX", "06HN33Z", "03HY32Z", ""), i10_pr4 = c("", "", "B54CZZA", 
"02HV33Z", ""), i10_pr5 = c("", "", "0W993ZX", "B548ZZA", "")), class = c("grouped_df", 
"tbl_df", "tbl", "data.frame"), row.names = c(NA, -5L), groups = structure(list(
    visitlink = c(2481935, 4605446, 7023336, 7287173, 7466851
    ), .rows = structure(list(3L, 4L, 2L, 5L, 1L), ptype = integer(0), class = c("vctrs_list_of", 
    "vctrs_vctr", "list"))), class = c("tbl_df", "tbl", "data.frame"
), row.names = c(NA, -5L), .drop = TRUE))

# df2: 待匹配的列表
structure(list(CODE = structure(1:7, levels = c("GZ58ZZZ", "3C1ZX8Z", "0BH17EZ", "HZ89ZZZ", "02HV33Z", "HZ99ZZZ", "XW03351"), class = "factor")), row.names = c(NA, 6L), class = "data.frame")

解决方案

方法1:使用tidyverse工具链

先处理df2的因子类型问题,再对df1逐行检查匹配情况:

library(tidyverse)

# 将df2的CODE转为字符型,避免因子匹配误差
match_codes <- as.character(df2$CODE)

df_result <- df1 %>%
  ungroup() %>% # 取消分组,不影响后续逻辑
  rowwise() %>%
  mutate(
    match = as.integer(
      any(c_across(starts_with("i10_pr")) %in% match_codes & c_across(starts_with("i10_pr")) != "")
    )
  ) %>%
  ungroup() # 可选,恢复非分组状态

# 查看结果
df_result

说明:

  • c_across(starts_with("i10_pr")):提取当前行所有以i10_pr开头的列值
  • %in% match_codes:检查值是否在匹配列表中
  • & c_across(...) != "":排除空字符串的干扰
  • any():判断当前行是否有任意值满足匹配条件
  • as.integer():将逻辑值转为1/0格式

方法2:使用base R实现

无需额外包,直接用基础函数完成:

# 处理匹配码的类型问题
match_codes <- as.character(df2$CODE)

# 筛选所有以i10_pr开头的列名
pr_cols <- grep("^i10_pr", names(df1), value = TRUE)

# 逐行检查匹配情况
df1$match <- apply(df1[pr_cols], 1, function(row) {
  as.integer(any(row %in% match_codes & row != ""))
})

# 查看结果
df1

说明:

  • grep("^i10_pr", names(df1), value = TRUE):精准筛选目标列
  • apply(..., 1, ...):对每一行执行匹配检查
  • 逻辑与tidyverse方法一致,确保只匹配非空的有效代码

内容的提问来源于stack exchange,提问作者ltong

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 19:54:28