You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于多模式向量为R数据框新增列并优先级赋值

解决方案

问题回顾

需要给数据集df新增列col2,按以下规则赋值:

  • 匹配matching1的行 → 1
  • 匹配matching2但未匹配matching1的行 → 2
  • 其余行 → 3

数据集与匹配向量定义:

# 数据集df
df <- structure(list(col1 = c("a b", "d e", "g f", "h j", "j k", "y z", 
"e f", "b c", "f g", "c d", "y z", "t u")), class = "data.frame", row.names = c(NA, 
-12L))

# 匹配向量
matching1 <- c("a b", "b c", "c d")
matching2 <- c("c d","e f","f g")

方法一:使用dplyr(简洁直观)

利用dplyr::mutate结合case_when,通过%in%直接判断元素归属,且case_when的顺序会自动保证优先级(先匹配的规则不会被后序规则覆盖):

library(dplyr)

df <- df %>%
  mutate(col2 = case_when(
    col1 %in% matching1 ~ "1",  # 优先匹配matching1
    col1 %in% matching2 ~ "2",  # 仅匹配matching2且未被上一条命中的行
    TRUE ~ "3"                  # 剩余所有行
  ))

方法二:基础R实现(无依赖)

通过分步赋值的方式,先初始化默认值,再按优先级覆盖:

# 初始化col2为默认值3
df$col2 <- "3"
# 给匹配matching1的行赋值1
df$col2[df$col1 %in% matching1] <- "1"
# 给仅匹配matching2的行赋值2(用setdiff排除matching1中的元素)
df$col2[df$col1 %in% setdiff(matching2, matching1)] <- "2"

说明

两种方法都无需逐个列出匹配项,适配匹配向量元素较多的场景,最终输出完全符合预期要求。

内容的提问来源于stack exchange,提问作者flxflks

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 03:45:36